Image processing method and device, electronic equipment, storage medium and program product
By providing reference images and segmentation mask images to train the image segmentation model, the adaptability problem of different segmentation tasks is solved, and the efficient adaptability of the general image segmentation model is achieved.
Patent Information
- Application Number
- CN202410331614.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-21
- Publication Date
- 2025-09-23
AI Technical Summary
Existing image segmentation models need to be retrained for different segmentation tasks, which is time-consuming and labor-intensive, and difficult to adapt to a variety of image segmentation tasks.
Provide a reference image and its corresponding segmentation mask image, train the image segmentation model, use colors to mark different image areas, and implement a universal image segmentation model that can adapt to various segmentation tasks.
This eliminates the need to retrain the model for different segmentation tasks, improves the versatility and adaptability of the image segmentation model, and simplifies the image segmentation processing flow.
Smart Images

Figure CN120689602A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision processing technology, and specifically to an image processing method and device, electronic equipment, computer-readable storage medium, and computer program product. Background Art
[0002] In the field of computer vision processing technology, image segmentation refers to the process of dividing an image into several specific regions with unique properties and extracting the target of interest. Image segmentation is an important research direction in the field of computer vision, and the image segmentation task is a fundamental and key problem.
[0003] Image segmentation tasks can be classified into various types, such as semantic segmentation, instance segmentation, panoptic segmentation, person segmentation, medical image segmentation, and aerial image segmentation. Despite significant progress in recent years in developing algorithms for various image segmentation tasks, current image segmentation models are still limited to specific tasks and data types. Whenever faced with a new segmentation task or data type, the model must be retrained and a large amount of labeled data must be prepared, which is extremely time-consuming and labor-intensive.
[0004] Therefore, how to provide a universal segmentation solution applicable to a variety of image segmentation tasks is a technical problem that needs to be solved by those skilled in the art. Summary of the Invention
[0005] To solve the above technical problems, embodiments of the present application provide an image processing method and apparatus, an electronic device, a computer-readable storage medium, and a computer program product.
[0006] One aspect of an embodiment of the present application provides an image processing method, which includes: obtaining a reference image, a segmentation mask image corresponding to the reference image, and an image to be processed; wherein the segmentation mask image corresponding to the reference image contains the color corresponding to each pixel in the reference image and is used to characterize the image segmentation task; inputting the reference image, the segmentation mask image corresponding to the reference image, and the image to be processed into an image segmentation model to obtain the segmentation mask image corresponding to the image to be processed output by the image segmentation model; wherein the image segmentation model is obtained by performing prediction training on the segmentation mask image of at least two image samples with the same object, and in the segmentation mask images corresponding to each image sample, the color of the same object is the same; according to the segmentation mask image corresponding to the image to be processed, image segmentation processing is performed on the image to be processed.
[0007] Another aspect of an embodiment of the present application provides an image processing device, which includes: an image acquisition module, configured to acquire a reference image, a segmentation mask image corresponding to the reference image, and an image to be processed; wherein the segmentation mask image corresponding to the reference image contains the color corresponding to each pixel in the reference image and is used to characterize the image segmentation task; an image prediction module, configured to input the reference image, the segmentation mask image corresponding to the reference image, and the image to be processed into an image segmentation model to obtain a segmentation mask image corresponding to the image to be processed output by the image segmentation model; wherein the image segmentation model is obtained by performing prediction training on the segmentation mask images of at least two image samples having the same object, and in the segmentation mask images corresponding to each image sample, the color of the same object is the same; an image segmentation module, configured to perform image segmentation processing on the image to be processed according to the segmentation mask image corresponding to the image to be processed.
[0008] Another aspect of an embodiment of the present application provides an electronic device, comprising: one or more processors; and a memory for storing one or more programs, which, when executed by the one or more processors, enables the electronic device to implement the image processing method described above.
[0009] Another aspect of an embodiment of the present application provides a computer-readable storage medium having computer-readable instructions stored thereon. When the computer-readable instructions are executed by a processor of a computer, the computer is caused to execute the image processing method described above.
[0010] Another aspect of an embodiment of the present application provides a computer program product, including a computer program, which implements the image processing method described above when executed by a processor.
[0011] In the technical solution proposed in the embodiments of the present application, on the one hand, a general image segmentation model is provided, and by providing a reference image and a segmentation mask image corresponding to the reference image, the image segmentation model is guided to determine a specific segmentation task, thereby avoiding the tedious process of training segmentation models for different segmentation tasks; on the other hand, the segmentation mask image corresponding to the reference image is colored, and the segmentation mask image represents different image areas by different colors. Due to the richness of color types, the image segmentation model provided by the present application can adapt to multi-category or multi-instance situations under multiple segmentation tasks, thereby making the image segmentation model provided by the present application highly universal.
[0012] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 is a flowchart of an image processing method shown in an exemplary embodiment of the present application;
[0014] Figure 2 is a flowchart of an image processing method shown in another exemplary embodiment of the present application;
[0015] Figure 3 is a schematic diagram of a colored mask image corresponding to an exemplary image sample illustrated in this application;
[0016] Figure 4 This is an exemplary image sample splicing diagram illustrated in this application;
[0017] Figure 5 is a flowchart of an image processing method shown in another exemplary embodiment of the present application;
[0018] Figure 6 is another exemplary image stitching diagram illustrated in this application;
[0019] Figure 7 is a structural diagram of an exemplary image segmentation model illustrated in this application;
[0020] Figure 8 This is a workflow diagram of an exemplary image segmentation model;
[0021] Figure 9 is a block diagram of an image processing apparatus shown in an exemplary embodiment of the present application;
[0022] Figure 10 A schematic diagram of the structure of a computer system suitable for implementing an electronic device according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0023] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. When the following description refers to the drawings, identical numerals in different figures represent identical or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0024] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically separate entities. That is, these functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0025] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, while others may be combined or partially combined. Therefore, the actual execution order may vary depending on the actual situation.
[0026] In this application, "plurality" refers to two or more. "And / or" describes the relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates that the related objects are in an "or" relationship.
[0027] The terms "first," "second," "third," and "fourth," etc., in the specification and claims of this application and the accompanying drawings are used to distinguish different objects, not to describe a specific order. The terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements, but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.
[0028] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0029] First of all, it should be noted that the technical solution provided in this application relates to the field of artificial intelligence (AI) technology.
[0030] Artificial intelligence (AI) is the theory, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also involves studying the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.
[0031] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, pre-trained models, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or foundational models, can, after fine-tuning, be widely applied to downstream tasks across various AI domains.
[0032] Artificial intelligence technologies mainly include computer vision (CV), speech processing (Speech Technology), natural language processing (NLP), and machine learning (ML). The technical solutions provided in this application mainly involve computer vision and machine learning.
[0033] Specifically, computer vision is the science of making machines "see." More specifically, it refers to the use of cameras and computers to replace the human eye in identifying, tracking, and measuring objects, and further processing the images to create images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, aiming to build artificial intelligence systems that can extract information from images or multidimensional data. Large model technology has brought significant changes to the development of computer vision technology. Pre-trained models in the field of vision, such as the Swin Transformer, ViT, V-MOE, and MAE, can be fine-tuned to quickly and widely apply to specific downstream tasks. Computer vision technology generally includes image processing, image recognition, image semantic understanding, image retrieval, video processing, video semantic understanding, video content / action recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and other technologies. It also includes common biometric recognition technologies such as facial recognition and fingerprint recognition.
[0034] Machine learning is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, and transfer learning. Pretrained models are the latest development in deep learning, integrating these techniques.
[0035] The technical solution of this application is to provide a universal image segmentation model for handling diverse segmentation tasks. It treats the image segmentation problem as a universal visual perception model and unifies different segmentation tasks through a contextual learning framework. When performing image segmentation, by providing a reference image and a segmentation mask image corresponding to the reference image, the image segmentation model can infer the segmentation task through the context, and then perform the corresponding segmentation task on a new image to obtain the corresponding segmentation result.
[0036] The technical solution provided by this application will be introduced in detail below.
[0037] First see Figure 1 , Figure 1 FIG. 4 is a flowchart of an image processing method shown in an exemplary embodiment of the present application.
[0038] The image processing method can be executed by terminal devices such as smart phones, computers, intelligent voice interaction devices, smart home appliances, vehicle terminals, aircraft, etc., or by a server, and this embodiment does not limit this.
[0039] When the image processing method is specifically executed by a server, the server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0040] Furthermore, when the image processing method is specifically executed by a server, some input data involved in the image processing method may be input by a terminal device connected to the server by wire or wireless communication. For example, the reference image, the segmentation mask image corresponding to the reference image, and the image to be processed are all uploaded to the server by the user through a computer. Figure 1 As shown, the exemplary image processing method includes S110-S130, which are described in detail as follows:
[0041] S110, obtaining a reference image, a segmentation mask image corresponding to the reference image, and an image to be processed; the segmentation mask image corresponding to the reference image includes a color corresponding to each pixel in the reference image and is used to represent the image segmentation task.
[0042] First, image segmentation is an important research area in computer vision. Image segmentation is the process of dividing an image into several specific regions with unique properties and extracting the objects of interest. Alternatively, it can be understood as the process of dividing an image into several regions with similar properties.
[0043] There are many types of image segmentation tasks, such as semantic segmentation, instance segmentation, panoramic segmentation, person segmentation, medical image segmentation, and aerial image segmentation. Semantic segmentation refers to the use of semantic expressions in images to segment images, that is, to divide the image into different semantic regions, such as road regions, building regions, and sky regions. Instance segmentation, based on semantic segmentation, distinguishes different instances of objects of the same category in an image, such as distinguishing each person or vehicle in an image. Panoramic segmentation combines semantic segmentation and instance segmentation to further provide semantic understanding of objects and background in an image and segment object instances. Person segmentation refers to the segmentation of people in an image. Medical image segmentation refers to the specific segmentation of medical images, such as separating different organs or lesion areas. Aerial image segmentation refers to the specific segmentation of aerial images, such as separating networked or patchy areas such as roads, rivers, and forests in aerial images.
[0044] It can be seen from this that the image segmentation task can be different segmentation tasks divided according to different segmentation granularity, or different segmentation tasks divided according to different application scenarios, or different segmentation tasks divided in other ways. This embodiment does not limit the specific type of segmentation task.
[0045] Segmentation masks are a key concept in image segmentation. They enable fine-grained segmentation of image regions by assigning classification labels to each pixel. For example, each pixel is assigned a label indicating whether it belongs to the foreground or background, or to a different object category. A segmentation mask image records the segmentation mask information for the corresponding image.
[0046] In this embodiment, in order to adapt the general image segmentation model to the needs of various segmentation tasks, the segmentation mask image uses different colors to classify and mark each pixel. In other words, each color represents a classification mark. Therefore, the colors contained in the segmentation mask image can be used to achieve image segmentation. It can also be understood more generally that each color in the segmentation mask image represents a property, and the image area of each color in the segmentation mask image is also an area with similar properties in the corresponding image. For ease of understanding, for example, assuming that the segmentation mask image B corresponding to image A contains a total of 5 color blocks, then the segmentation mask image B can be used to divide image A into the corresponding 5 blocks.
[0047] Since the image segmentation model is universally applicable to different image segmentation tasks, when the image segmentation model is actually used, it is necessary to make the image segmentation model aware of the specific segmentation task in order to achieve correct image segmentation processing. Therefore, this embodiment needs to provide a reference image and a segmentation mask image of the reference image. The segmentation mask image corresponding to the reference image contains the color corresponding to each pixel in the reference image, which can be used to characterize the segmentation task corresponding to the image to be processed. In this way, the image segmentation model is guided to obtain the segmentation task based on the reference image, and then the corresponding image segmentation processing is performed on the image to be processed according to this segmentation task.
[0048] It should be noted that the image segmentation model is trained by predicting segmentation mask images for at least two image samples that share the same object. In the segmentation mask images corresponding to each image sample, the same object has the same color. Therefore, during the application phase of the image segmentation model, the reference image, the segmentation mask image corresponding to the reference image, and the image to be processed are input into the image segmentation model. This ensures that the coloring of the segmentation mask image for the image to be processed, as predicted by the image segmentation model, is consistent with the coloring of the segmentation mask image corresponding to the reference image. Therefore, for the image segmentation model, the segmentation mask image corresponding to the reference image can represent the image segmentation task for the image to be processed.
[0049] The different image segmentation task types in the above examples are used to illustrate how the color of each pixel contained in the segmentation mask image corresponding to the reference image can be used to characterize the image segmentation task. For example, in the scenario of a person segmentation task, the image area containing the person is marked as color 1 in the segmentation mask image corresponding to the reference image. In the segmentation mask image corresponding to the image to be processed output by the image segmentation model, the image area containing the person is also marked as color 1, so that when the person segmentation is subsequently performed, the image area with color 1 can be directly segmented out. Alternatively, the image areas corresponding to different people are marked with different colors in the segmentation mask image corresponding to the reference image. In the segmentation mask image corresponding to the image to be processed output by the image segmentation model, compared with the segmentation mask image corresponding to the reference image, the color of the image area corresponding to the same person should be consistent, while the color of the image areas corresponding to different people should be different. It should be noted that the image areas that do not contain people in the segmentation mask image can be marked with other colors or not marked with colors, and this is not limited here.
[0050] For example, in the scenario of semantic segmentation tasks, image areas with different semantics are marked with different colors in the segmentation mask image corresponding to the reference image, such as the road area is marked as color 3, the building area is marked as color 4, and the sky area is marked as color 5. In the segmentation mask image corresponding to the image to be processed output by the image segmentation model, the color of the image area with the same semantics as the reference image is consistent with the color in the segmentation mask image corresponding to the reference image, while the color of the image area with different semantics from the reference image can be other random colors, which facilitates the subsequent execution of specific image segmentation processing according to different colors.
[0051] It should be understood that in different image segmentation scenarios, based on the different image segmentation tasks, the subsequent image segmentation processing process specifically performed based on the segmentation mask image output by the image segmentation model should also be different, and this embodiment does not limit this.
[0052] S120 , inputting the reference image and the segmentation mask image corresponding to the reference image into the image segmentation model to obtain the segmentation mask image corresponding to the image to be processed output by the image segmentation model.
[0053] As described above, the image segmentation model determines the segmentation task based on the reference image and the segmentation mask image corresponding to the reference image, and thus predicts the segmentation mask image corresponding to the image to be processed based on the segmentation task. Therefore, there is no need to train different segmentation models for different image segmentation tasks. Instead, the general image segmentation model provided in this application is used to predict the segmentation mask image of the image to be processed under the guidance of the reference image, so that the image segmentation processing of the image to be processed can be implemented subsequently based on the predicted segmentation mask image of the image to be processed.
[0054] S130 , performing image segmentation processing on the image to be processed according to the segmentation mask image corresponding to the image to be processed.
[0055] The present application performs segmentation processing on the image to be processed based on the segmentation mask image corresponding to the image to be processed, which can be understood as a process of dividing the image to be processed into different areas according to the segmentation mask image corresponding to the image to be processed.
[0056] The specific image segmentation processing process is usually related to the actual application scenario, and this embodiment does not elaborate on this process in detail. For example, for a portrait segmentation task, the image segmentation processing of the image to be processed according to the segmentation mask image corresponding to the image to be processed can be to retain the portrait area and delete the non-portrait area. For another example, for a medical image segmentation task, the image segmentation processing of the image to be processed according to the segmentation mask image corresponding to the image to be processed can be to segment out the target area in the medical image, and the segmented target area can be used for quantitative analysis of tissue volume, diagnosis, localization of pathologically changed tissue, depiction of anatomical structure, treatment planning, etc.
[0057] From the above, it can be seen that in the technical solution proposed in the embodiment, on the one hand, a general image segmentation model is provided, and by providing a reference image and a segmentation mask image corresponding to the reference image, the image segmentation model is guided to determine a specific segmentation task, thereby avoiding the tedious process of training segmentation models for different segmentation tasks; on the other hand, the segmentation mask image corresponding to the reference image is colored, and the segmentation mask image uses different colors to represent image areas of different properties. Due to the richness of color types, the image segmentation model provided by this application can adapt to multi-category or multi-instance situations under multiple segmentation tasks, thereby making the image segmentation model provided by this application have high versatility.
[0058] It should also be noted that the technical solution proposed in this embodiment can be applied to the field of transportation.
[0059] The Intelligent Traffic System (ITS), also known as the Intelligent Transportation System (ITS), effectively integrates advanced science and technology (information technology, computer technology, data communication technology, sensor technology, electronic control technology, automatic control theory, operations research, artificial intelligence, etc.) into transportation, service control, and vehicle manufacturing. This strengthens the connection between vehicles, roads, and users, thereby forming an integrated transportation system that ensures safety, improves efficiency, enhances the environment, and conserves energy. Intelligent Vehicle Infrastructure Cooperative Systems (IVICS), also known as VIS, are a development direction of ITS. VIS utilizes advanced wireless communications and next-generation internet technologies to implement dynamic, real-time information exchange between vehicles and roads. Based on the collection and integration of dynamic traffic information across all time and space, it conducts active vehicle safety control and road collaborative management, fully realizing effective coordination between people, vehicles, and roads, ensuring traffic safety, and improving traffic efficiency, thereby forming a safe, efficient, and environmentally friendly road transportation system.
[0060] The technical solution proposed in this embodiment can be specifically applied to the recognition of dynamic traffic information in all time and space within a vehicle-road cooperative system, thereby achieving effective collaboration between people, vehicles, and roads. For example, the image processing solution proposed in this embodiment can segment images captured by intelligent vehicles or roadside equipment, and effectively identify traffic conditions based on the segmentation results, thereby implementing active vehicle safety control and road cooperative management based on the identified traffic conditions.
[0061] Of course, it should be noted that the technical solution proposed in this embodiment can also be applied to other fields, and this embodiment does not limit this.
[0062] In another exemplary embodiment, Figure 2 As shown in Figure 1 The embodiment shown further includes S210-S240, which are described in detail as follows:
[0063] S210 , obtaining at least two image samples having the same object from a training dataset.
[0064] This embodiment discloses the training process of the image segmentation model.
[0065] To train a universal segmentation model, the training process first randomly selects at least two image samples from the training dataset that share the same object. At least two image samples sharing the same object can be understood as containing image content of the same category, or containing image content belonging to the same instance of the object. For example, if two images both contain a pillow, then the pillow is the same object in both images. For another example, if two images both contain multiple portraits of a person, and the multiple portraits contain the same person, then the same person is also the same object in both images.
[0066] S220 , performing coloring processing on the segmentation mask images corresponding to at least two image samples to obtain a colored mask image corresponding to each image sample. The color of the same object in each colored mask image is the same.
[0067] In this embodiment, the segmentation mask images corresponding to at least two image samples are colored using the same coloring method in order to color the same objects in at least two image samples with the same color, so that the image segmentation model has the ability to recognize contextually consistent information after training, and accordingly obtains the colored mask images corresponding to each image sample.
[0068] Figure 3 FIG is a schematic diagram of a color mask image corresponding to an exemplary image sample illustrated in this application. Figure 3 As shown, the two exemplary image samples on the left are image sample 1 and image sample 2, in which the image areas corresponding to the same person are colored in the same color; the two exemplary image samples on the right are image sample 3 and image sample 4, in which the pillow areas are colored in the same color.
[0069] S230 , stitching at least two image samples to obtain a stitched image sample, and stitching the coloring mask images corresponding to the at least two image samples to obtain a stitched coloring mask image.
[0070] Figure 4 This is an exemplary image sample splicing diagram illustrated in the present application, which specifically illustrates splicing image sample A and image sample B to obtain a spliced image sample, and splicing a colored mask image corresponding to image sample A and a colored mask image corresponding to image sample B to obtain a spliced colored mask image.
[0071] It can be seen that by splicing at least two image samples and splicing the coloring mask images corresponding to the at least two image samples, the obtained spliced image samples and the spliced coloring mask images can be made to correspond to each other in content.
[0072] S240 , masking a portion of the image area in the spliced color mask image, and training the image segmentation model based on the masked color mask image, the spliced image samples, and the spliced color mask image.
[0073] During the model training phase, this embodiment randomly blocks the spliced colored mask image, meaning that certain image regions are randomly masked. The image segmentation model is trained by predicting the colors of these blocked regions and comparing the predicted colors with the unblocked colored mask image for loss supervision. This transforms the image segmentation model training process into a solution to the contextual colorization problem, using color mapping to enable the model to rely on contextual information for segmentation task discrimination.
[0074] It should also be noted that, in another exemplary embodiment, in the process of coloring the segmentation mask images corresponding to at least two image samples using the same coloring method, random coloring is used for coloring. Random coloring can be understood as, in each round of iterative training of the image segmentation model, the color assigned to the image content with consistent context is random. Random coloring can also be understood as, in at least two iterations of model training, the color of the same object in the colored mask image is different. Figure 3 The coloring mask image shown on the right is an example. If this iteration randomly colors the pillow area Figure 3 In the blue shown, in subsequent iterations, the pillow area may be randomly colored yellow or other colors. This shows that this embodiment treats various types of image segmentation tasks as a random coloring problem. The trained image segmentation model does not rely on specific color information to perform segmentation task discrimination, but instead uses contextual information for reasoning. This allows it to better adapt and perform well when faced with new segmentation tasks and data, thereby improving the generalization ability of the image segmentation model.
[0075] Figure 5 FIG. 1 is a flowchart of an image processing method shown in another exemplary embodiment of the present application. Figure 5 As shown in Figure 1 The embodiment shown further includes S510-S520, which are described in detail as follows:
[0076] S510 , stitching the reference image and the image to be processed to obtain a stitched reference image, and stitching the segmentation mask image corresponding to the reference image and the full occlusion mask image corresponding to the image to be processed to obtain a stitched mask image.
[0077] It can be understood that in this embodiment, during the inference stage of the image segmentation model, with reference to the training stage of the image segmentation model, a reference image and a segmentation mask image corresponding to the reference image are given as reference images, and then a to-be-processed image that is to perform the same segmentation task as the reference image is given. The segmentation mask image corresponding to the to-be-processed image is equivalent to being completely blocked, so that the image segmentation model predicts the color information of all the blocked areas, thereby obtaining the segmentation mask image corresponding to the to-be-processed image.
[0078] In the training phase of the reference image segmentation model, it is necessary to splice the reference image with the image to be processed to obtain a spliced reference image, and to splice the segmentation mask image corresponding to the reference image with the full occlusion mask image corresponding to the image to be processed to obtain a spliced mask image. It can be understood that the full occlusion mask image corresponding to the image to be processed can refer to Figure 6 As shown, the entire image area of the full occlusion mask image is covered in black, while the segmentation mask image corresponding to the reference image contains color information of different area blocks.
[0079] S520 , inputting the spliced reference image and the spliced mask image into an image segmentation model, so that the image segmentation model outputs a segmentation mask image corresponding to the image to be processed, and the segmentation mask image corresponding to the image to be processed contains the predicted color information.
[0080] By inputting the spliced reference image and the spliced mask image into the image segmentation model, the image segmentation model learns the specific segmentation task based on the segmentation mask image corresponding to the reference image. Based on this segmentation task, the model then predicts the color information corresponding to the fully occluded mask image corresponding to the image to be processed, thereby obtaining the segmentation mask image corresponding to the image to be processed. As a result, the segmentation mask image corresponding to the image to be processed output by the image segmentation model contains the color information predicted by the model.
[0081] In order to make the prediction result of the image segmentation model more accurate, at least two reference images can be used, so at least two reference images need to be integrated.
[0082] As an exemplary embodiment, at least two reference images can be spliced together, and the size of the spliced reference image can be reset to match the size of a separate reference image to obtain a resized reference image, so that the resized reference image can be used as a reference image for splicing with the image to be processed; and, the segmentation mask images corresponding to at least two reference images can be spliced together, and the size of the spliced mask image can be reset to match the size of a separate segmentation mask image to obtain a resized segmentation mask image, so that the resized segmentation mask image can be used as a segmentation mask image for splicing with the fully occluded mask image.
[0083] In summary, this implementation involves stitching together multiple reference images and resizing them to a size suitable for the image segmentation model. The segmentation mask images corresponding to each of the multiple reference images also require similar processing. However, this implementation compresses the reference image information to a certain extent, resulting in a loss of image resolution.
[0084] As another exemplary embodiment, each of at least two reference images is spliced with the image to be processed to obtain at least two spliced reference images; and the segmentation mask image corresponding to each reference image is spliced with the fully occluded mask image corresponding to the image to be processed to obtain at least two spliced mask images. By inputting the at least two spliced reference images and the two spliced mask images into an image segmentation model, the image segmentation model predicts the color information corresponding to the fully occluded mask image corresponding to the image to be processed, thereby obtaining a segmentation mask image corresponding to the image to be processed. It can be seen that this embodiment does not cause a loss in image resolution. For the image segmentation model, in the case of a single reference image, for example, the model input feature dimensions are (2, 896, 448, 3), in this embodiment, the model input feature dimensions are (n, 896, 448, 3), but this does not affect the feature dimensions of the model's final output. It can be understood that n represents image data, 896 represents image height, 448 represents image width, and 3 represents the number of image channels.
[0085] The model structure of the image segmentation model is introduced below.
[0086] like Figure 7 As shown, the image segmentation model includes a first fully connected network layer, an encoding layer, and a decoding layer connected in sequence.
[0087] The first fully connected network layer is used to map the spliced reference image and the spliced mask image into image features of the target feature dimension respectively. The target feature dimension is understood to be the feature dimension adapted by the image segmentation model. For example, it is first necessary to use a preset size of unit pixel block to divide the spliced reference image and the spliced mask image into blocks. For example, if the size of the unit pixel block is 16*16, the spliced reference image and the spliced mask image will be divided into 1568 (56*28) pixel blocks respectively. By stretching each pixel block 16*16*3 into a 768-dimensional feature, and then mapping it to 1024 dimensions through the first fully connected layer, the resulting feature dimension is (2, 1568, 1024).
[0088] The encoding layer encodes the image features of the spliced reference image and the spliced mask image to obtain encoded features. Exemplarily, the encoding layer consists of multiple encoding blocks connected in sequence. The encoding blocks can be, for example, transformer blocks. As can be understood, each transformer block consists of two main components: an attention mechanism and a feedforward component. The attention mechanism is used to handle context.
[0089] In some exemplary embodiments, the multiple coding blocks included in the coding layer are divided into two parts, namely a first coding block part and a second coding block part. During the encoding process, the first coding block part first extracts feature maps from the image features of the target feature dimension of each of the spliced reference image and the spliced mask image; then, the feature maps of the spliced reference image and the spliced mask image are fused to obtain a fused feature map; and then, the second coding block part extracts features from the fused feature map to obtain coding features.
[0090] It should be noted that since each input image consists of two images and a corresponding segmentation mask image, the input resolution is equivalent to twice that of the traditional non-universal segmentation model training process, which results in high computational costs. To address this issue, this embodiment introduces a feature fusion operation during the encoding process, which can significantly reduce the amount of computation required for the entire encoding process, thereby saving computational costs. It can also be understood that this embodiment reduces computational costs by introducing a feature fusion operation during the encoding process to merge the early features of the input image and the segmentation mask image.
[0091] For example, assuming that the coding layer adopts an architecture in which 24 transformer blocks are connected in sequence, the feature maps can be extracted from the image features of the target feature dimensions of the spliced reference image and the spliced mask image respectively through the first X transformer blocks. Feature fusion is performed after passing through X transformer blocks. The feature dimension can be reduced through feature fusion, and then the coding features are further extracted for the fused features through the remaining Y transformer blocks. Compared with the method of directly using 24 transformer blocks to extract coding features for the spliced reference image and the spliced mask image respectively, the method of introducing the feature fusion operation in the encoding process mentioned in this embodiment can obviously save computing costs and will not cause performance degradation.
[0092] It should be noted that in this example, the sum of X and Y should be 24, and both X and Y are positive integers. For example, when X is 10 and Y is 14, it means that the first 10 transformer blocks extract feature maps from the image features of the target feature dimension of the spliced reference image and the spliced mask image respectively, and then perform feature fusion. The remaining 14 transformer blocks continue to extract coding features based on the fused features.
[0093] In another exemplary embodiment, by setting the number of coding blocks contained in the second coding block part to be greater than the number of coding blocks contained in the first coding block part, the node for merging the early features of the input image and the segmentation mask image is brought to the front, thereby further reducing the computational cost.
[0094] For example, assuming that the coding layer adopts an architecture in which 24 transformer blocks are connected in sequence, the first three transformer blocks can be used to extract feature maps from the image features of the target feature dimensions of the spliced reference image and the spliced mask image. Feature fusion is performed after the three transformer blocks. Feature fusion can reduce the feature dimension, thereby saving computing costs and not causing performance degradation. For example, even in the case of multiple reference images, if the feature dimension of the input image of the model is (2, 1568, 1024), this embodiment can reduce the feature dimension to (1, 1568, 1024). If the feature dimension of the input image of the model is (n, 1568, 1024), the feature dimension can still be reduced to (1, 1568, 1024). Afterwards, the fused features will be input into the remaining 21 transformer blocks for encoding to obtain the encoded features.
[0095] Compared with the first 10 transformer blocks that extract feature maps from the image features of the target feature dimensions of the spliced reference image and the spliced mask image respectively, and then perform feature fusion, and continue to extract encoding features for the fused features through the remaining 14 transformer blocks, this example moves the feature fusion node forward to the third transformer block, thereby bringing the node for merging the early features of the input image and the segmentation mask image forward, which can further reduce the computational cost.
[0096] In other exemplary embodiments, considering that in the case of a multi-coding structure, the image represented by the output features of the coding blocks closer to the front is clearer, and the image represented by the output features of the coding blocks closer to the back is more blurred, however, it is difficult to determine whether the clearer image features are more conducive to image segmentation or the more blurred image features are more conducive to image segmentation. Therefore, this embodiment splices the output features of at least two coding blocks contained in the second coding block part to obtain a multi-scale feature, and uses this multi-scale feature as the final coding feature. That is, the output features of at least two coding blocks used for feature extraction of the fusion feature map are spliced to obtain a multi-scale feature. For example, the output features of the coding blocks of the 5th, 11th, 17th, and 23rd layers can be connected in series to form a multi-scale feature with a feature dimension of (1, 1568, 4096). It should be noted that this embodiment does not limit the multiple coding blocks used to form the multi-scale feature.
[0097] The decoding layer is used to decode the coded features to obtain the color information corresponding to the predicted fully occluded mask image. Exemplarily, the decoding layer is composed of a second fully connected layer and a convolutional network layer connected in sequence. The second fully connected layer is used to map the coded features output by the coding layer to pixel-level features. The dimension of the input features of the second fully connected layer is (1, 1568, 4096), and the dimension of the output features is (1, 896, 448, 128). The convolutional network layer is used to perform mask prediction processing on the pixel-level features output by the second fully connected layer. The feature dimension of the resulting mask prediction is (1, 896, 448, 3), thereby obtaining the color information corresponding to the fully occluded mask image corresponding to the image to be processed.
[0098] Figure 8 This is a workflow diagram of an exemplary image segmentation model. Figure 8 It can be seen that the feature dimensions of the spliced reference image and the spliced mask image are (1, 896, 448, 3), so that the feature dimension of the input first connection network layer is (2, 896, 448, 3), and the feature dimension of the output after processing by the first fully connected network layer is (2, 1568, 1024). The dimensions of the feature maps extracted by the three transformer blocks are (1, 1568, 1024), and the feature dimensions input to the remaining 21 transformer blocks after feature fusion are (1, 1568, 1024). The encoded features obtained by encoding are multi-scale features with a dimension of (1, 1568, 4096). The multi-scale features are processed by the decoding layer to obtain features with a dimension of (1, 896, 448, 128). The dimension of the segmentation mask map corresponding to the fully occluded image is (1, 896, 448, 3).
[0099] As can be seen from the above, in the image segmentation solution provided in this application, a universal image segmentation model is proposed, which can handle a variety of segmentation tasks, such as semantic segmentation, instance segmentation, foreground segmentation, character segmentation, medical image segmentation, aerial image segmentation, etc. Through a unified model structure and training method, the tedious process of retraining the segmentation model for different segmentation tasks and data types is avoided. In practical applications, by giving a reference image and a corresponding segmentation mask image, various segmentation tasks can be easily performed without retraining the model or a large amount of labeled data. This convenience and applicability has very important application value in the field of computer vision. In addition, the simple and effective model structure proposed in this embodiment also makes the image segmentation model easy to train and deploy, and can run efficiently on different hardware platforms.
[0100] Figure 9 FIG. 1 is a block diagram of an image processing apparatus according to an exemplary embodiment of the present invention. Figure 9 As shown, the exemplary image processing apparatus includes:
[0101] An image acquisition module 910 is configured to acquire a reference image, a segmentation mask image corresponding to the reference image, and an image to be processed; wherein the segmentation mask image corresponding to the reference image includes a color corresponding to each pixel in the reference image and is used to represent the image segmentation task;
[0102] The image prediction module 920 is configured to input the reference image, the segmentation mask image corresponding to the reference image, and the image to be processed into an image segmentation model to obtain a segmentation mask image corresponding to the image to be processed output by the image segmentation model; wherein the image segmentation model is obtained by performing segmentation mask image prediction training on at least two image samples having the same object, and in the segmentation mask images corresponding to the respective image samples, the same object has the same color;
[0103] The image segmentation module 930 is configured to perform image segmentation processing on the image to be processed according to the segmentation mask image corresponding to the image to be processed.
[0104] In another exemplary embodiment, the image processing apparatus further includes a model training module, and the model training module is configured to:
[0105] Get at least two image samples of the same object from the training dataset;
[0106] Coloring the segmentation mask images corresponding to at least two image samples using the same coloring method to obtain colored mask images corresponding to each image sample;
[0107] splicing at least two image samples to obtain a spliced image sample, and splicing the colored mask images corresponding to the at least two image samples to obtain a spliced colored mask image;
[0108] Part of the image area in the spliced color mask image is blocked, and the image segmentation model is trained based on the blocked color mask image, the spliced image samples, and the spliced color mask image.
[0109] In another exemplary embodiment, in at least two iterations of model training, the colors of the same object in the colored mask image are different.
[0110] In another exemplary embodiment, the image prediction module 920 includes a splicing processing module, which is configured to:
[0111] Splicing the reference image with the image to be processed to obtain a spliced reference image, and splicing the segmentation mask image corresponding to the reference image with the full occlusion mask image corresponding to the image to be processed to obtain a spliced mask image;
[0112] The spliced reference image and the spliced mask image are input into the image segmentation model, so that the image segmentation model outputs a segmentation mask image corresponding to the image to be processed, and the segmentation mask image corresponding to the image to be processed contains the predicted color information.
[0113] In another exemplary embodiment, the number of reference images is at least two, and the image prediction module 920 further includes a pre-processing module, which is configured to:
[0114] splicing at least two reference images, and resizing the size of the spliced reference image to match the size of the individual reference images to obtain a resized reference image, and using the resized reference image as a reference image for splicing with the image to be processed;
[0115] Furthermore, the segmentation mask images corresponding to at least two reference images are stitched together, and the size of the stitched mask image is resized to match the size of the individual segmentation mask images to obtain a resized segmentation mask image, so as to use the resized segmentation mask image as the segmentation mask image for stitching with the full occlusion mask image.
[0116] In another exemplary embodiment, the number of reference images is at least two, and the image prediction module 920 is further configured to:
[0117] splicing each of the at least two reference images with the image to be processed to obtain at least two spliced reference images;
[0118] Furthermore, the segmentation mask image corresponding to each reference image is spliced with the full occlusion mask image corresponding to the image to be processed to obtain at least two spliced mask images.
[0119] In another exemplary embodiment, the image segmentation model includes a first fully connected network layer, an encoding layer, and a decoding layer connected in sequence; the image prediction module 920 is further configured to:
[0120] Mapping the spliced reference image and the spliced mask image into image features of the target feature dimension respectively through the first fully connected network layer;
[0121] Performing feature encoding processing on the image features of the spliced reference image and the spliced mask image through the encoding layer to obtain encoding features;
[0122] The encoded features are decoded through the decoding layer to obtain the segmentation mask image corresponding to the image to be processed.
[0123] In another exemplary embodiment, the image prediction module 920 is further configured to:
[0124] Extracting feature maps from the image features of the target feature dimensions of the spliced reference image and the spliced mask image respectively;
[0125] Fusing the feature maps of the spliced reference image and the spliced mask image to obtain a fused feature map;
[0126] Feature extraction is performed on the fused feature map to obtain coding features; wherein the number of coding blocks contained in the second coding block part is greater than the number of coding blocks contained in the first coding block part.
[0127] In another exemplary embodiment, the coding layer is composed of a plurality of coding blocks connected in sequence; the prediction module 920 is further configured to:
[0128] Feature maps are extracted from image features of target feature dimensions of the spliced reference image and the spliced mask image respectively through the first coding block portion, and feature extraction is performed on the fused feature map through the second coding block portion to obtain coding features; wherein the number of coding blocks contained in the second coding block portion is greater than the number of coding blocks contained in the first coding block portion.
[0129] In another exemplary embodiment, the coding layer is composed of a plurality of coding blocks connected in sequence; the prediction module 920 is further configured to: concatenate the output features of at least two coding blocks for feature extraction of the fusion feature map to obtain multi-scale features;
[0130] Use multi-scale features as encoding features.
[0131] In another exemplary embodiment, the decoding layer is composed of a second fully connected layer and a convolutional network layer connected in sequence; the prediction module 920 is further configured to:
[0132] The encoded features are mapped to pixel-level features through the second fully connected network layer;
[0133] The pixel-level features are processed by mask prediction through the convolutional network layer to obtain the color information corresponding to the fully occluded mask image.
[0134] It should be noted that the image processing devices provided in the above embodiments and the image processing methods provided in the above embodiments are based on the same concept. The specific manner in which the various modules and units perform their operations has been described in detail in the method embodiments and will not be repeated here. In actual applications, the image processing devices provided in the above embodiments can, as needed, allocate the aforementioned functions to different functional modules, i.e., divide the internal structure of the device into different functional modules to perform all or part of the functions described above. This is not a limitation herein.
[0135] The image processing device provided in the above embodiment, on the one hand, guides the general image segmentation model to determine a specific segmentation task by providing a reference image and a segmentation mask image corresponding to the reference image, thereby avoiding the tedious process of training segmentation models for different segmentation tasks; on the other hand, the segmentation mask image corresponding to the reference image is colored, and the segmentation mask image uses different colors to represent image areas of different properties. Due to the richness of color types, the image segmentation model provided by this application can adapt to multi-category or multi-instance situations under multiple segmentation tasks, thereby making the image segmentation model provided by this application highly universal.
[0136] An embodiment of the present application also provides an electronic device, comprising: one or more processors; and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the electronic device implements the image processing method provided in the above-mentioned embodiments.
[0137] Figure 10 The following is a schematic diagram showing the structure of a computer system suitable for implementing an electronic device according to an embodiment of the present application. Figure 10 The computer system 1000 of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0138] like Figure 10As shown, the computer system 1000 includes a central processing unit (CPU) 1001, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage part 1008 into the random access memory (RAM) 1003, such as executing the method described in the above embodiment. Various programs and data required for system operation are also stored in the RAM 1003. The CPU 1001, ROM 1002 and RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0139] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, and the like; an output section 1007 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 1008 including a hard disk and the like; and a communication section 1009 including a network interface card such as a LAN (Local Area Network) card or a modem. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as needed. Removable media 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 1010 as needed, so that computer programs read therefrom can be installed into the storage section 1008 as needed.
[0140] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1009, and / or installed from a removable medium 1011. When the computer program is executed by the central processing unit (CPU) 1001, the various functions defined in the system of the present application are executed.
[0141] It should be noted that the computer-readable medium shown in the embodiment of the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. The computer program contained in the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0142] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. Among them, each box in the flowchart or block diagram can represent a module, program segment, or part of the code, and the above-mentioned module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0143] The units involved in the embodiments described in this application may be implemented by software or hardware, and the units described may also be set in a processor. In some cases, the names of these units do not constitute limitations on the units themselves.
[0144] Another aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the image processing method described above. The computer-readable storage medium may be included in the electronic device described in the above embodiments, or may exist independently and not be incorporated into the electronic device.
[0145] Another aspect of the present application provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the image processing method provided in each of the above embodiments.
[0146] The above content is only a preferred exemplary embodiment of the present application and is not intended to limit the implementation scheme of the present application. Ordinary technicians in this field can easily make corresponding changes or modifications based on the main ideas and spirit of the present application. Therefore, the scope of protection of the present application shall be based on the scope of protection required by the claims.
[0147] It is understandable that in the specific implementation of this application, any image and other related data are involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.
Claims
1. An image processing method, characterized in that: The method comprises: Obtaining a reference image, a segmentation mask image corresponding to the reference image, and an image to be processed; wherein the segmentation mask image corresponding to the reference image includes a color corresponding to each pixel in the reference image and is used to represent the image segmentation task; Inputting the reference image, the segmentation mask image corresponding to the reference image, and the image to be processed into an image segmentation model to obtain a segmentation mask image corresponding to the image to be processed output by the image segmentation model; wherein the image segmentation model is obtained by performing segmentation mask image prediction training on at least two image samples having the same object, and in the segmentation mask images corresponding to the respective image samples, the same object has the same color; Perform image segmentation processing on the image to be processed according to the segmentation mask image corresponding to the image to be processed.
2. The method according to claim 1, characterized in that The method further comprises: Get at least two image samples of the same object from the training dataset; performing coloring processing on the segmentation mask images corresponding to the at least two image samples to obtain a colored mask image corresponding to each image sample, wherein the color of the same object in each colored mask image is the same; splicing the at least two image samples to obtain a spliced image sample, and splicing the colored mask images corresponding to the at least two image samples to obtain a spliced colored mask image; Part of the image area in the spliced color mask image is blocked, and the image segmentation model is trained based on the blocked color mask image, the spliced image samples, and the spliced color mask image.
3. The method according to claim 2, characterized in that The same object has different colors in the colorized mask image within at least two iterations of model training.
4. The method according to claim 1, wherein The step of inputting the reference image, the segmentation mask image corresponding to the reference image, and the image to be processed into an image segmentation model to obtain the segmentation mask image corresponding to the image to be processed output by the image segmentation model includes: splicing the reference image with the image to be processed to obtain a spliced reference image, and splicing the segmentation mask image corresponding to the reference image with the full occlusion mask image corresponding to the image to be processed to obtain a spliced mask image; The spliced reference image and the spliced mask image are input into the image segmentation model, so that the image segmentation model outputs a segmentation mask image corresponding to the image to be processed, and the segmentation mask image corresponding to the image to be processed contains predicted color information.
5. The method according to claim 4, characterized in that The number of the reference images is at least two; and the method further includes: splicing the at least two reference images, and resizing the size of the spliced reference image to match the size of the individual reference images to obtain a resized reference image, and using the resized reference image as a reference image for splicing with the image to be processed; Furthermore, the segmentation mask images corresponding to the at least two reference images are spliced together, and the size of the spliced mask image is resized to match the size of the individual segmentation mask images to obtain a resized segmentation mask image, so as to use the resized segmentation mask image as the segmentation mask image for splicing with the fully occluded mask image.
6. The method according to claim 4, characterized in that The number of the reference images is at least two; the step of splicing the reference images with the image to be processed to obtain a spliced reference image, and splicing the segmentation mask image corresponding to the reference image with the full occlusion mask image corresponding to the image to be processed to obtain a spliced mask image, includes: splicing each of the at least two reference images with the image to be processed to obtain at least two spliced reference images; Furthermore, the segmentation mask image corresponding to each reference image is respectively spliced with the full occlusion mask image corresponding to the image to be processed to obtain at least two spliced mask images.
7. The method according to claim 4, characterized in that The image segmentation model includes a first fully connected network layer, an encoding layer, and a decoding layer connected in sequence; inputting the spliced reference image and the spliced mask image into the image segmentation model so that the image segmentation model outputs a segmentation mask image corresponding to the image to be processed, comprising: Mapping the spliced reference image and the spliced mask image into image features of target feature dimensions respectively through the first fully connected network layer; Performing feature encoding processing on image features of the spliced reference image and the spliced mask image through the encoding layer to obtain encoding features; The encoding features are decoded by the decoding layer to obtain a segmentation mask image corresponding to the image to be processed.
8. The method according to claim 7, characterized in that The performing feature encoding processing on the image features of the spliced reference image and the spliced mask image through the encoding layer to obtain encoding features includes: Extracting feature maps from image features of target feature dimensions of the stitched reference image and the stitched mask image respectively; Fusing the feature maps of the spliced reference image and the spliced mask image to obtain a fused feature map; Feature extraction is performed on the fused feature map to obtain coding features.
9. The method according to claim 8, characterized in that The coding layer is composed of a plurality of coding blocks connected in sequence; a first coding block portion extracts feature maps from image features of target feature dimensions of the spliced reference image and the spliced mask image, and a second coding block portion performs feature extraction on the fused feature map to obtain coding features; The number of coding blocks contained in the second coding block part is greater than the number of coding blocks contained in the first coding block part.
10. The method according to claim 8, characterized in that The coding layer is composed of a plurality of coding blocks connected in sequence; and the feature extraction of the fused feature map to obtain coding features includes: splicing output features of at least two encoding blocks used for feature extraction of the fused feature map to obtain multi-scale features; The multi-scale features are used as the encoding features.
11. The method according to claim 7, characterized in that The decoding layer is composed of a second fully connected layer and a convolutional network layer connected in sequence; the decoding layer decodes the encoded features to obtain the predicted color information corresponding to the fully occluded mask image, including: Mapping the encoded features into pixel-level features through the second fully connected network layer; The convolutional network layer performs mask prediction processing on the pixel-level features to obtain color information corresponding to the fully occluded mask image.
12. An image processing device, characterized in that: The device comprises: an image acquisition module configured to acquire a reference image, a segmentation mask image corresponding to the reference image, and an image to be processed; wherein the segmentation mask image corresponding to the reference image includes a color corresponding to each pixel in the reference image and is used to represent the image segmentation task; an image prediction module configured to input the reference image, a segmentation mask image corresponding to the reference image, and the image to be processed into an image segmentation model to obtain a segmentation mask image corresponding to the image to be processed output by the image segmentation model; wherein the image segmentation model is obtained by performing segmentation mask image prediction training on at least two image samples having the same object, and in the segmentation mask images corresponding to the respective image samples, the same object has the same color; The image segmentation module is configured to perform image segmentation processing on the image to be processed according to the segmentation mask image corresponding to the image to be processed.
13. An electronic device, characterized in that: include: one or more processors; The memory is used to store one or more programs, and when the one or more programs are executed by the one or more processors, the electronic device implements the method according to any one of claims 1 to 10.
14. A computer-readable storage medium, characterized in that Computer-readable instructions are stored thereon, and when the computer-readable instructions are executed by a processor of a computer, the computer is caused to execute the method according to any one of claims 1 to 11.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.