Image processing method and device, equipment and storage medium
By extracting the target area in image processing for feature extraction and segmentation, the uncertainty is estimated based on the feature information of the target area map, and the uncertainty problem of the influence of the auxiliary segmentation model is solved, and the accuracy of the image segmentation result is improved.
Patent Information
- Application Number
- CN202410242759.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-01
- Publication Date
- 2025-09-02
AI Technical Summary
In the prior art, the uncertainty of the segmentation result of the image segmentation model is affected by the accuracy of the auxiliary segmentation model, resulting in inaccuracy problems.
By extracting the target area in the image, feature extraction and image segmentation are performed, uncertainty is estimated based on the feature information of the target area map, and dependence on auxiliary segmentation models is avoided.
The accuracy of uncertainty estimation of image segmentation results is improved, the error caused by inaccuracy of auxiliary segmentation models is reduced, and the reliability of image processing is improved.
Smart Images

Figure CN120580464A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular to an image processing method, apparatus, device, and storage medium. Background Art
[0002] Currently, in fields such as biometrics and virtual reality, it is necessary to segment eye images and use the segmentation results to handle subsequent tasks. However, since the segmentation results may be inaccurate, it is necessary to estimate the uncertainty of the segmentation results to assist in the processing of subsequent tasks.
[0003] In related technologies, an image segmentation model is used to directly segment the eye image to be segmented, and the segmentation results of the image segmentation model are used as the final segmentation result. In addition, an auxiliary segmentation model is used to segment the eye image to be segmented again. The difference between the segmentation results of the auxiliary segmentation model and the image segmentation model is used to determine the uncertainty of the segmentation result of the image segmentation model.
[0004] Since the accuracy of the segmentation result of the auxiliary segmentation model is not high enough, when the segmentation result of the auxiliary segmentation model is inaccurate, the uncertainty of the segmentation result of the image segmentation model will also be inaccurate. Summary of the Invention
[0005] The embodiments of the present application provide an image processing method, apparatus, device, and storage medium. The technical solutions provided by the embodiments of the present application are as follows:
[0006] According to one aspect of an embodiment of the present application, there is provided an image processing method, the method comprising:
[0007] Extracting a target area in the image to obtain a target area map, wherein pixels of the target area respectively belong to one of a plurality of preset pixel categories;
[0008] Performing feature extraction on the target area map to obtain feature information of the target area map;
[0009] Performing image segmentation on the target area map based on the feature information to obtain an image segmentation result of the target area map, wherein the image segmentation result indicates the preset pixel category to which each pixel in the target area map belongs;
[0010] Obtaining, based on the feature information, a confidence level of the preset pixel category to which each pixel in the target area map belongs;
[0011] The uncertainty of the image segmentation result is obtained according to the confidence of the preset pixel category to which each pixel in the target area map belongs.
[0012] According to one aspect of an embodiment of the present application, a model training method is provided, the method comprising:
[0013] Extracting a target area in the first sample image to obtain a first target area map, wherein pixels of the target area respectively belong to one of a plurality of preset pixel categories;
[0014] Inputting the first target area map into a feature extraction network of a trained image segmentation model, and having the feature extraction network of the trained image segmentation model output feature information of the first target area map;
[0015] Inputting the feature information into a feature classification network of the trained image segmentation model, and having the feature classification network of the trained image segmentation model output an image segmentation result of the first target area image, wherein the image segmentation result indicates the preset pixel category to which each pixel in the first target area image belongs;
[0016] Inputting the feature information into a confidence estimation model, obtaining the confidence of the preset pixel category to which each pixel in the first target area image belongs from the confidence estimation model, and using the confidence of the preset pixel category to which each pixel in the first target area image belongs to obtain the uncertainty of the image segmentation result;
[0017] According to the feature information and the confidence of the preset pixel category to which each pixel in the first target area map belongs, the parameters of the confidence estimation model are adjusted to obtain a trained confidence estimation model.
[0018] According to one aspect of an embodiment of the present application, there is provided an image processing apparatus, the apparatus comprising:
[0019] A region extraction module is used to extract a target region in an image and obtain a target region map, wherein pixels of the target region respectively belong to one of a plurality of preset pixel categories;
[0020] A feature extraction module is used to extract features from the target area map to obtain feature information of the target area map;
[0021] an image segmentation module, configured to perform image segmentation on the target area map based on the feature information to obtain an image segmentation result of the target area map, wherein the image segmentation result indicates the preset pixel category to which each pixel in the target area map belongs;
[0022] A confidence determination module, configured to obtain the confidence of the preset pixel category to which each pixel in the target area map belongs based on the feature information;
[0023] The uncertainty estimation module is used to obtain the uncertainty of the image segmentation result according to the confidence of the preset pixel category to which each pixel in the target area map belongs.
[0024] According to one aspect of an embodiment of the present application, a model training device is provided, the device comprising:
[0025] a target region extraction module, configured to extract a target region from the first sample image to obtain a first target region map, wherein pixels of the target region respectively belong to one of a plurality of preset pixel categories;
[0026] A feature information extraction module, configured to input the first target area map into a feature extraction network of a trained image segmentation model, and output feature information of the first target area map from the feature extraction network of the trained image segmentation model;
[0027] an image segmentation processing module, configured to input the feature information into a feature classification network of the trained image segmentation model, and output an image segmentation result of the first target area image from the feature classification network of the trained image segmentation model, wherein the image segmentation result indicates the preset pixel category to which each pixel in the first target area image belongs;
[0028] a confidence information acquisition module, configured to input the feature information into a confidence estimation model, and obtain, from the confidence estimation model, the confidence of the preset pixel category to which each pixel in the first target area image belongs, wherein the confidence of the preset pixel category to which each pixel in the first target area image belongs is used to obtain the uncertainty of the image segmentation result;
[0029] The estimation model adjustment module is used to adjust the parameters of the confidence estimation model according to the feature information and the confidence of the preset pixel category to which each pixel in the first target area map belongs, so as to obtain a trained confidence estimation model.
[0030] According to one aspect of an embodiment of the present application, a computer device is provided, comprising a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the above-mentioned image processing method or model training method.
[0031] According to one aspect of an embodiment of the present application, a computer-readable storage medium is provided, in which a computer program is stored. The computer program is loaded and executed by a processor to implement the above-mentioned image processing method or model training method.
[0032] According to one aspect of an embodiment of the present application, a computer program product is provided, which includes a computer program, and the computer program is loaded and executed by a processor to implement the above-mentioned image processing method or model training method.
[0033] The technical solutions provided by the embodiments of the present application include at least the following beneficial effects:
[0034] In this application, the uncertainty estimation of the image segmentation results does not require the use of the segmentation results of another auxiliary segmentation model, but instead performs uncertainty estimation based on the feature information of the target area map, thereby avoiding the problem of poor accuracy of uncertainty estimation caused by inaccurate segmentation results of the auxiliary segmentation model, and improving the accuracy of uncertainty estimation.
[0035] Furthermore, by extracting the target region from the image, a target region map (which can be understood as a sub-image of the image) is obtained. Based on the feature information of the target region map, a confidence assessment is performed on each pixel in the target region map. Then, based on the confidence of each pixel, the uncertainty of the image segmentation result is obtained. Because the target region map is a sub-image of the image, it is obtained by removing most of the non-target areas in the image to include the target region. Therefore, it can remove interference from other image content in the image besides the target region, improve the accuracy of the pixel confidence estimation, and thus improve the accuracy of the uncertainty estimation of the image segmentation result. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 This is a schematic diagram of an implementation environment for a solution provided by an embodiment of the present application;
[0037] Figure 2 is a flowchart of an image processing method provided by an embodiment of the present application;
[0038] Figure 3 is a distribution diagram of uncertainty of an image segmentation result provided by an embodiment of the present application;
[0039] Figure 4 This is a flowchart of a model training method provided by one embodiment of the present application;
[0040] Figure 5 This is a flowchart of target detection model training provided by one embodiment of the present application;
[0041] Figure 6 This is a flowchart of image segmentation model training provided by one embodiment of the present application;
[0042] Figure 7 This is a flowchart of confidence estimation model training provided by one embodiment of the present application;
[0043] Figure 8 This is a flowchart of the image processing method and model training method provided by one embodiment of the present application;
[0044] Figure 9 This is a flowchart of the use of the model after deployment provided by an embodiment of the present application;
[0045] Figure 10 is a block diagram of an image processing device provided by one embodiment of the present application;
[0046] Figure 11 This is a block diagram of a model training device provided by one embodiment of the present application;
[0047] Figure 12 This is a structural block diagram of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0048] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0049] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also involves studying the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.
[0050] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, pre-trained models, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or basic models, can be fine-tuned and widely applied to downstream tasks across various AI disciplines. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0051] Computer vision (CV) is the science of making machines "see." Specifically, it refers to using cameras and computers to replace the human eye in identifying and measuring objects, and then performing further image processing to create images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, aiming to build artificial intelligence systems that can extract information from images or multidimensional data. Large model technology has brought significant changes to the development of computer vision technology. Pre-trained models in the field of vision, such as the swin-transformer, ViT, V-MOE, and MAE, can be fine-tuned to quickly and widely apply to specific downstream tasks. Computer vision technology generally includes image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / action recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and other technologies. It also includes common biometric recognition technologies such as facial recognition and fingerprint recognition.
[0052] Machine Learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning by demonstration. Pretrained models are the latest development in deep learning, integrating these techniques.
[0053] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, autonomous driving, drones, digital twins, virtual humans, robots, artificial intelligence generated content (Artificial Intelligence Generated Content, referred to as AIGC), conversational interaction, smart medical care, smart customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0054] The solutions provided in the embodiments of this application involve technologies such as computer vision and machine learning based on artificial intelligence, which are specifically introduced and explained through the following embodiments.
[0055] Please refer to Figure 1 , which shows a schematic diagram of a solution implementation environment provided by an embodiment of the present application. The solution implementation environment may include: a model training device 10 and a model use device 20.
[0056] The model training device 10 is an electronic device with data calculation, processing and storage functions. The model training device 10 can be either a terminal device or a server. The model training device 10 can include but is not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, game consoles, wearable devices, multimedia playback devices, augmented reality (AR) devices, virtual reality (VR) devices and other electronic devices. The model training device 10 is used to train image segmentation models and confidence estimation models. Optionally, it is also used to train target detection models.
[0057] In the present application, the image segmentation model can be a neural network model, and the confidence estimation model can also be a neural network model. For example, the neural network model can be a model constructed based on a convolutional neural network. Optionally, the model training device 10 can use a machine learning method to train the above-mentioned image segmentation model and confidence estimation model, so that the trained image segmentation model has better image segmentation performance, and the trained confidence estimation model has better confidence estimation performance.
[0058] In some embodiments, based on the trained image segmentation model, the training process of the confidence estimation model is as follows (this is only a brief description here, and the specific training process can be found in the following embodiments): extract the target area from the first sample image 13 to obtain a first target area map 14; input the first target area map 14 into the feature extraction network of the trained image segmentation model to obtain the feature information of the first target area map 14; input the feature information of the first target area map 14 into the feature classification network of the trained image segmentation model to obtain the corresponding image segmentation result; input the feature information of the first target area map 14 into the confidence estimation model, and output the confidence of the preset pixel category to which each pixel in the first target area map 14 belongs; based on the confidence and feature information of the preset pixel category to which each pixel in the first target area map 14 belongs, calculate the loss function value of the confidence estimation model; adjust the parameters of the confidence estimation model according to the loss function value to obtain the trained confidence estimation model.
[0059] In some embodiments, the training process of the image segmentation model is as follows (this is only a brief description, and the specific training process is described in the following embodiments): extract the target area from the second sample image 11 to obtain the second target area Figure 12 ; The second target area Figure 12 Input into the feature extraction network of the image segmentation model to obtain the second target area Figure 12 The second target area Figure 12 The feature information is input into the feature classification network of the image segmentation model to obtain the corresponding image segmentation result; based on the image segmentation result and the labeled image segmentation result, the loss function value of the image segmentation model is calculated; according to the loss function value, the parameters of the image segmentation model are adjusted to obtain the trained image segmentation model.
[0060] The trained image segmentation model and confidence estimation model are used in the image processing method.
[0061] The model-using device 20 is an electronic device with data calculation, processing, and storage functions. The model-using device 20 can be either a terminal device or a server. The model-using device 20 can include, but is not limited to, mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, game consoles, wearable devices, multimedia playback devices, augmented reality devices, virtual reality devices, cloud technology platforms, intelligent robots, smart transportation terminal systems, and driving control systems. The model-using device 20 uses the trained image segmentation model and the trained confidence estimation model to process the image and obtain the image segmentation result corresponding to the image and the uncertainty of the image segmentation result.
[0062] The model training device 10 and the model using device 20 can be two independent devices or the same device.
[0063] For example, the server mentioned above can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms, but is not limited to these.
[0064] The embodiments of the present application can be applied to various scenarios. For example, based on different target areas, they can be divided into eye area application scenarios, hand area application scenarios, head area application scenarios, road area application scenarios, etc.
[0065] In the eye area application scenario, the embodiments of the present application can be applied to image processing tasks such as biometrics, line of sight estimation, and eye image quality assessment.
[0066] For example, in the biometric image processing task of the eye area application scenario, after image processing is performed on the image containing the eye area, a target area map containing the eye area (also called an eye area map) is obtained, and image segmentation is performed on the target area map to obtain an image segmentation result, and then the iris area is determined from the image segmentation result. Since the iris texture contains a large number of detailed features, and the detailed features remain unchanged during the life cycle of the organism, the identity of the organism can be identified by analyzing the iris area.
[0067] For example, in the line of sight estimation image processing task of the eye area application scenario, after image processing is performed on the image containing the eye area, a target area map containing the eye area (also called an eye area map) is obtained, and image segmentation is performed on the target area map to obtain an image segmentation result, and then the pupil area is determined from the image segmentation result. Since the radius of the iris of the eye and its position on the face are fixed, the line of sight direction of the organism can be estimated by analyzing the pupil area and combining it with other features.
[0068] For example, in the eye image quality assessment of the eye area application scenario, after image processing is performed on the image containing the eye area, a target area map containing the eye area (also referred to as an eye area map) is obtained, and image segmentation is performed on the target area map to obtain an image segmentation result. The pupil area, iris area and white of the eye area are determined from the image segmentation result, and the quality of the target area map is evaluated by combining the above three areas.
[0069] In the hand area application scenario, the embodiments of the present application can be applied to image processing tasks such as biometrics and hand posture estimation.
[0070] For example, in the biometric image processing task in the hand area application scenario, after image processing is performed on the image containing the hand area, a target area map containing the hand area (also called a hand area map) is obtained, and image segmentation is performed on the target area map to obtain an image segmentation result, and then the palm area is determined from the image segmentation result. Since each person's palm print features are unique and the palm print features remain unchanged during the life cycle of the organism, the biometric identity can be identified by analyzing the palm area.
[0071] For example, in the hand posture recognition image processing task in the hand area application scenario, after image processing is performed on the image containing the hand area, a target area map containing the hand area (also called a hand area map) is obtained, and image segmentation is performed on the target area map to obtain an image segmentation result. Then, each finger area and palm area is determined from the image segmentation result, and the shape, size and position of the above-mentioned finger area are analyzed to estimate the hand posture.
[0072] In the head area application scenario, the embodiments of the present application can be applied to image processing tasks such as biometrics and head posture estimation.
[0073] For example, in the biometric image processing task of the head area application scenario, after image processing is performed on the image containing the head area, a target area map containing the head area (also called a head area map) is obtained, and image segmentation is performed on the target area map to obtain an image segmentation result. The facial area is then determined from the image segmentation result, and features of the facial area are extracted, analyzed and matched to identify the biometric identity.
[0074] For example, in the head posture estimation image processing task of the head area application scenario, after image processing is performed on the image containing the head area, a target area map containing the head area (also called a head area map) is obtained, and image segmentation is performed on the target area map to obtain an image segmentation result. The head area and facial area are then determined from the image segmentation result. The head posture can be estimated by analyzing the position, angle, shape of the head and the position of the face in the head area.
[0075] In road area application scenarios, the embodiments of the present application can be applied to image processing tasks such as road condition analysis and traffic control.
[0076] For example, in the road condition analysis image processing task of the road area application scenario, the image containing the driving road can be processed to obtain a target area map containing the road area (also called a road area map), and the target area map is segmented to obtain an image segmentation result. The area where the road object is located is determined from the image segmentation result, and the shape, size, etc. of the road object are analyzed. The current road condition can be analyzed to assist the driver in driving safely.
[0077] For example, in the traffic control image processing task of the road area application scenario, the image containing the driving road can be processed to obtain a target area map containing the road area (also called a road area map), and the target area map is segmented to obtain an image segmentation result. The road area and vehicle area are then extracted from the image segmentation result. Based on the number of vehicles in the road area, the current road traffic situation is judged, and traffic control is performed based on the current road traffic situation.
[0078] Of course, the application scenarios and examples of image segmentation of target area maps introduced above are merely exemplary and explanatory. The technical solution of this application can be applied to any application scenario in which one or more areas are segmented from an image containing a detection target, and this application does not limit this.
[0079] Please refer to Figure 2 , which shows a flow chart of an image processing method provided by an embodiment of the present application. The execution subject of each step of the method can be a computer device, for example, the computer device can be Figure 1 The model in the solution implementation environment shown uses a device 20. The method may include at least one of the following steps (210-250):
[0080] Step 210 : extracting a target area from the image to obtain a target area map, wherein pixels in the target area respectively belong to one of a plurality of preset pixel categories.
[0081] Images can be in a variety of formats and representations, including but not limited to the following formats or representations: Joint Photographic Experts Group (JPEG), Portable Network Graphics (PNG), Bitmap (BMP), Tagged Image File Format (RAW), Scalable Vector Graphics (SVG), and other image formats. Each pixel uses a numerical value to represent attributes such as brightness, color, and position on the image. For example, in a color image, three channels, red, green, and blue (RGB), are used to represent color. By controlling the intensity values of the three channels, the colors in the image can be accurately specified and adjusted. In black and white or grayscale images, a single channel is used to represent brightness or grayscale levels. The resolution of an image refers to the number of pixels in each direction of the image. The higher the resolution, the clearer the image.
[0082] In some embodiments, the image may be a video frame. For example, the image may be a video frame in a video. In some embodiments, the image may also be a single picture.
[0083] In general, the image in step 210 can be any image with a limited number of pixel categories, or it can be understood that all pixels in the image belong to a limited number of preset pixel categories.
[0084] In the eye region application scenario, an image refers to an image containing the eye region. Optionally, the eye region can be a human eye or the eye of another creature, such as a cat, dog, or penguin. The eye region in the image can include at least one of the following structures: eyeball, iris, pupil, sclera, etc. The iris is the annular portion located between the pupil and sclera, and its surface contains radiating blood vessels, forming an iris texture.
[0085] In the hand region application scenario, an image refers to an image containing a hand region. Optionally, the hand can be a human hand or a hand of another creature, such as a gorilla or monkey. The hand in the image can include at least one of the following structures: wrist, fingers, and palm. Fingers and palms have complex and unique patterns.
[0086] In the head region application scenario, an image refers to an image containing a head region. Optionally, the head region can be a human head or a head of another creature, such as a chicken, cat, or dog. The head region in the image can include at least one of the following structures: head, neck, or face.
[0087] In a road area scenario, the image may refer to an image including a road area. Optionally, the road may include road objects such as vehicles, humans, animals, and stationary objects.
[0088] The target region refers to the local area of the image that contains the detection target. The target region map can be considered a sub-image of the image. It can contain only the target region, or it can contain the target region and the non-target region surrounding the target region. For example, it can be a minimum rectangular image that includes the target region and the non-target region surrounding the target region. The detection target refers to the target object in the image that needs to be identified and located. For example, the detection target can be an eye, hand, head, road, etc.
[0089] In some embodiments, the non-target area in the target area map refers to an area that does not contain the detection target. For example, in an eye area application scenario, the non-target area may include the following elements: glasses, hair, eyelashes, facial skin, etc. For example, in a hand area application scenario, the non-target area may include the following elements: bracelets, watches, rings, etc. For example, in a head area application scenario, the non-target area may include the following elements: hairpins, hats, other body parts, etc. For example, in a road area application scenario, the non-target area includes trees on both sides of the road, street lights, etc.
[0090] In some embodiments, the preset pixel categories refer to pre-set pixel categories. The preset pixel categories may vary depending on the application scenario, as described in the embodiments below. The preset pixel categories to which pixels in the target area belong are multiple known preset pixel categories. It can be understood that the names and number of preset pixel categories to which all pixels in the target area belong are known, but the preset pixel category to which each pixel belongs is unknown.
[0091] In some embodiments, based on the target area, a minimum rectangular frame including the target area is determined in the image, and then the image content within the minimum rectangular frame is captured from the image to obtain a target area map.
[0092] In some embodiments, an image may include one or more target regions. For example, when an image includes only one target region, a single target region map is obtained when the target region is extracted from the image. For example, when an image includes multiple target regions, multiple target region maps are obtained when the target regions are extracted from the image.
[0093] In some embodiments, the target area map can be extracted from the image by manual interception, or the target area map can be identified and extracted from the image by an AI model.
[0094] In some embodiments, an image is input into a target detection model, which outputs location information of a target area in the image; based on the location information, the target area is captured from the image to obtain a target area map.
[0095] In some embodiments, the minimum rectangular frame including the target area is determined based on the location information of the target area, and then the image content within the minimum rectangular frame is captured from the image to obtain a target area map.
[0096] The target detection model is a neural network model. For example, the target detection model can be a model built based on a convolutional neural network (CNN), which is used to obtain the location information of the target area from the image.
[0097] In some embodiments, an image is input into an object detection model to extract spatial features of the image. Exemplarily, after inputting the image into the object detection model, the spatial features of the image are obtained, and the spatial features include features such as texture, shape, edges, local structure, and spatial distribution of the image.
[0098] In some embodiments, the object detection model processes the extracted spatial features of the image and outputs location information of the target region in the image. For example, the object detection model may perform operations such as pooling and regression on the extracted spatial features of the image to locate the target region in the image and obtain the location information of the target region.
[0099] In some embodiments, the position information is used to indicate the position of the target area in the image. The position information may include at least one of the following: coordinate information of the minimum rectangular frame including the target area in the image, and size information of the minimum rectangular frame including the target area.
[0100] Exemplarily, when the position information includes coordinate information of the upper left vertex of the minimum rectangular frame including the target area and size information of the minimum rectangular frame including the target area, the minimum rectangular frame including the target area can be determined in the image.
[0101] Exemplarily, when the position information includes coordinate information of four vertices of the minimum rectangular frame including the target area, the minimum rectangular frame including the target area can also be determined in the image.
[0102] Through the above method, the target area is located in the image through the target detection model, and the corresponding target area map is captured from the image, realizing the automatic detection and extraction of the target area using AI means, which helps to improve the efficiency of the entire image processing process.
[0103] Step 220: extract features from the target area map to obtain feature information of the target area map.
[0104] In some embodiments, the feature information is a feature map extracted using a neural network. A feature map is a quantitative representation of features extracted from the target region map. It can be represented using structured data such as a sparse matrix, a sparse vector, or a dense vector. The specific feature representation method used is determined by relevant technical personnel and is not limited in this application. The size of the feature map is determined by the size of the target region map and the structure of the neural network.
[0105] For example, assuming that the above-mentioned neural network is a convolutional neural network, including 9 convolution kernels, and the size of the target area map is 1080×1080 pixels, the size of the feature map extracted by the convolutional neural network is 1080×1080×9, and each pixel in the target area map corresponds to 9 features.
[0106] In some embodiments, the above-mentioned neural network is a neural network for extracting feature information of the target area map, which can be a convolutional neural network, a fully convolutional neural network, a feedforward neural network, or a neural network including an attention mechanism, etc., which can be used to extract feature information of the image. This application does not limit this.
[0107] In related technologies, directly using neural networks to extract features from images requires a large amount of computing resources, such as high-performance processors and memory, and places high demands on the application environment of the neural network.
[0108] In the above manner, by only performing feature extraction on the target area map and not on other parts of the image, most useless feature extraction is reduced, the efficiency of feature extraction is improved, and the waste of computing resources is reduced.
[0109] Step 230 : performing image segmentation on the target region map based on the feature information to obtain an image segmentation result of the target region map. The image segmentation result indicates the preset pixel category to which each pixel in the target region map belongs.
[0110] Image segmentation is the process of identifying the target region in the target area from the target region map and distinguishing the target region from the non-target region. The target region is the region in the target region map that contains only the target region.
[0111] In some embodiments, feature information is input into a feature classification network of an image segmentation model, and the feature classification network outputs an image segmentation result of an image region map, wherein the image segmentation model includes a feature extraction network and a feature classification network, and the feature extraction network is used to extract feature information; the feature information is input into a confidence estimation model, and the uncertainty of the image segmentation result is obtained based on the output features of the confidence estimation model.
[0112] The feature classification network is used to classify each pixel in the target region map based on the feature information of the target region map. The feature classification network is a type of neural network. For example, the feature classification network can be a convolutional neural network, a fully connected neural network (FCN), a deep convolutional neural network, etc.
[0113] In some embodiments, the multiple preset pixel categories and the number of pixels in the target area are determined. The preset pixel category to which each pixel in the target area belongs is determined during subsequent image segmentation. For example, a person skilled in the art can assign the multiple preset pixel categories to the pixels in the target area using preset pixel category names or numbers.
[0114] In some embodiments, the preset pixel categories to which each of the above-mentioned pixels belongs include: recognition object area and non-recognition object area. For example, in the biometric image processing task of the eye area application scenario, the recognition object is the iris, then the preset pixel category to which each pixel belongs is any one of the following: iris area, non-iris area. For example, in the biometric image processing task of the head area application scenario, the recognition object is the face, then the preset pixel category to which each pixel belongs is any one of the following: face area, non-face area. For example, in the biometric image processing task of the hand area application scenario, the recognition object is the palm, then the preset pixel category to which each pixel belongs is any one of the following: palm area, non-palm area. For example, in the road condition analysis image processing task of the road area application scenario, the recognition object is a road object, then the preset pixel category to which each pixel belongs is any one of the following: road object area, non-road object area.
[0115] For example, if the size of the target area map is 920×1080 pixels, the size of the image segmentation result corresponding to the target area map is also 920×1080. Assuming that the image segmentation result is represented by a binary mask, the value of each element in the image segmentation result is 0 or 1, that is, it indicates whether the pixel in the target area map corresponding to each element in the image segmentation result belongs to or does not belong to the recognition object area. For example, 0 represents belonging to the recognition object area, and 1 represents not belonging to the recognition object area, or, 0 represents not belonging to the recognition object area, and 1 represents not belonging to the recognition object area.
[0116] In some embodiments, the size of the image segmentation result is equal to the size of the target region map, and each element in the image segmentation result corresponds to whether a pixel in the target region map belongs to the identified object region. Optionally, the image segmentation result can be represented using a binary mask or other digital encoding, which is not limited in this application.
[0117] In some embodiments, the identification target may include multiple objects.
[0118] The preset pixel category to which each pixel belongs refers to the preset pixel category assigned to each pixel during feature classification, and the preset pixel category is pre-set by relevant technical personnel.
[0119] In some embodiments, the preset pixel category for classifying each pixel is determined based on processing requirements.
[0120] In some embodiments, when there is a need to process only a single identification object, when performing feature classification on each pixel of the target area map, the preset pixel category to which each pixel belongs can be set to only two preset pixel categories: identification object and non-identification object, that is, the target area map is divided into the area belonging to the identification object and the area not belonging to the identification object.
[0121] For example, in a pupil segmentation scenario, if there is a need to process the pupil area, when performing feature classification on each pixel of the eye area map, the preset pixel category to which each pixel belongs can be set to two preset pixel categories: pupil area and non-pupil area, that is, the eye area map is divided into areas belonging to the pupil and areas that do not belong to the pupil.
[0122] In some embodiments, when there is a need to process multiple types of identification objects, when performing feature classification on each pixel of the target area map, the preset pixel category to which each pixel belongs can be set to the above-mentioned multiple preset pixel categories and non-identification objects, that is, the target area map is divided into identification object areas belonging to at least one preset pixel category among the multiple preset pixel categories and non-identification object areas.
[0123] For example, if there are processing requirements for the pupil area, iris area and eye area, when performing feature classification on each pixel of the eye area map, the preset pixel category to which each pixel belongs can be set to four preset pixel categories: pupil area, iris area, eye area, and non-eye area, that is, the eye area map is divided into an area belonging to the pupil, an area belonging to the iris, an area belonging to the eye, and an area that does not belong to the pupil, iris, or eye.
[0124] For other preset pixel category combinations, the application process is similar and will not be described in detail here.
[0125] Through the above method, relevant technicians can segment the target area map according to needs and obtain image segmentation results suitable for different application scenarios.
[0126] Step 240 : Obtain the confidence level of the preset pixel category to which each pixel in the target area map belongs based on the feature information.
[0127] Confidence is a measure of how reliably a pixel belongs to a certain category. A higher confidence level indicates a more reliable result; a lower confidence level indicates a less reliable result.
[0128] In some embodiments, feature information is input into a confidence estimation model, which outputs confidence information. The confidence information includes a confidence feature matrix of each pixel in the target area map. The confidence feature matrix of the pixel is used to determine the confidence of the preset pixel category to which the pixel belongs. Based on the confidence feature matrix of each pixel in the target area map, the confidence of the preset pixel category to which each pixel in the target area map belongs is obtained.
[0129] In some embodiments, the confidence estimation model is used to estimate the confidence of each pixel in the target region map. The confidence estimation model can be a neural network model, for example, a model constructed by a convolutional neural network, a fully connected neural network, or the like.
[0130] In some embodiments, for each pixel in the target region map, the trace of the confidence feature matrix of the pixel is determined as the confidence of the preset pixel category to which the pixel belongs.
[0131] In some embodiments, the confidence feature matrix corresponding to a pixel is a diagonal matrix. A diagonal matrix is a square matrix in which all non-main diagonal elements are zero. The trace of a diagonal matrix is equal to the sum of the elements on the main diagonal. In other words, the trace of the confidence feature matrix for a pixel can be obtained by adding the elements on the main diagonal of the confidence feature matrix.
[0132] For example, if the size of the feature information is 100×200×D, then after inputting the feature information into the confidence estimation model, the size of the confidence information obtained is also 100×200×D×D. Based on the above example, for each pixel in the target area map, the corresponding confidence feature matrix size is D×D, and all elements except those on the main diagonal are zero.
[0133] In the above manner, by calculating the confidence level corresponding to each pixel, the reliability of each pixel belonging to the corresponding preset pixel category is determined, which can better measure the quality of the image segmentation result.
[0134] Step 250 : Obtain the uncertainty of the image segmentation result according to the confidence of the preset pixel category to which each pixel in the target area map belongs.
[0135] In this application, uncertainty is used to indicate the accuracy of the image segmentation result. The higher the uncertainty, the less accurate the image segmentation result; the lower the uncertainty, the more accurate the image segmentation result.
[0136] For example, Figure 3 As shown, it shows a distribution diagram of the uncertainty of the image segmentation result provided by an embodiment of the present application, Figure 3Sub-figures (a), (b), and (c) of the iris segmentation dataset are applied to the iris segmentation scenario and are the distribution diagrams of the uncertainty of the iris segmentation results obtained using the LPW dataset, the Dikablis dataset, and the OpenEDS dataset, respectively. Sub-figure (a) shows the distribution curve 31 of the uncertainty of the iris segmentation results obtained using the LPW dataset, sub-figure (b) shows the distribution curve 32 of the uncertainty of the iris segmentation results obtained using the Dikablis dataset, and sub-figure (c) shows the distribution curve 33 of the uncertainty of the iris segmentation results obtained using the OpenEDS dataset. It can be seen that the uncertainty distribution of the iris segmentation results obtained from different datasets basically follows a Gaussian distribution. Therefore, the uncertainty of the iris segmentation results can be used to measure the accuracy of the iris segmentation results.
[0137] In some embodiments, the confidence level of the image segmentation result may also be obtained based on the confidence level of the preset pixel category to which each pixel in the target region map belongs.
[0138] In some embodiments, the uncertainty of the image segmentation result is obtained according to the sum of the logarithms of the confidences of the preset pixel categories to which each pixel in the target area map belongs.
[0139] For example, assuming that x is any pixel in the target area map, the uncertainty formula of the image segmentation result can be expressed as:
[0140]
[0141] Among them, s out Represents the uncertainty of the image segmentation result, Σ x Represents the confidence feature matrix of pixel x, tr(*) represents the trace of the matrix, that is, the confidence feature matrix Σ of pixel x x The first ∑ on the right side of the equation represents the sum of the operation results corresponding to each pixel in the target area map.
[0142] Through the above method, a method for determining the uncertainty of the image segmentation result is provided, which comprehensively considers the confidence of the preset pixel category to which each pixel in the target area map belongs to determine the uncertainty of the image segmentation result, which helps to improve the accuracy of the determined uncertainty.
[0143] After obtaining the image segmentation result and its corresponding uncertainty, it can be determined whether to output the image segmentation result in the following manner.
[0144] In some embodiments, if the uncertainty is less than a preset threshold, the image segmentation result is determined to be valid; or, if the uncertainty is greater than or equal to the preset threshold, the image segmentation result is determined to be invalid.
[0145] In some embodiments, the above-mentioned preset threshold is a value used to compare with the uncertainty of the image segmentation result. It can be pre-set based on experience, experimental data, or other factors. The specific value is set by relevant technical personnel according to needs and is not limited by this application. For example, if the uncertainty range is 0 to 1, the preset threshold can be a value between 0 and 1, such as 0.4 or 0.5, which is not limited by this application.
[0146] Whether the image segmentation result is valid can be understood as whether the accuracy of the image segmentation result or the segmentation quality meets the requirements of relevant technicians in the image segmentation task. In other words, when the uncertainty of the image segmentation result is less than the above-mentioned preset threshold, it means that the segmentation quality of the image segmentation result is good and meets the requirements of relevant technicians in the image segmentation task; when the uncertainty of the image segmentation result is greater than or equal to the above-mentioned preset threshold, it means that the segmentation quality of the image segmentation result is poor and does not meet the requirements of relevant technicians in the image segmentation task, and it is discarded.
[0147] For example, assume that the uncertainty of the image segmentation result is s out , the above process can be expressed by the following formula:
[0148]
[0149] Among them, R out Indicates whether the image segmentation result is valid, and th represents the preset threshold used to compare with the uncertainty of the image segmentation result.
[0150] By using the above method, the quality of the image segmentation result obtained is high, and by setting different preset thresholds, image segmentation results suitable for different needs can be obtained.
[0151] After obtaining the above-mentioned effective image segmentation results, the image segmentation results can be subsequently processed or analyzed in combination with actual business needs.
[0152] In some embodiments, when the uncertainty is less than a preset threshold, subsequent image processing tasks are performed based on the image segmentation results. In other words, when the image segmentation results are valid, the image segmentation results can be further processed according to business needs.
[0153] In some embodiments, a designated area is segmented from the image segmentation results, and the designated area is analyzed and processed to obtain an analysis result. For example, in a biometric image processing task in an eye area application scenario, the iris area is segmented from the image segmentation results, and detailed features in the iris area are compared and analyzed to determine the user's identity.
[0154] Through the above method, when the uncertainty of the image segmentation result is less than the preset threshold, the user can post-process the image segmentation result according to different business needs, which helps to meet the diverse needs of users.
[0155] In some embodiments, the detection target is the eye, and the target area map is the eye area map; the iris area in the eye area map is determined according to the image segmentation result, and identity recognition is performed based on the iris area; or, the pupil area in the eye area map is determined according to the image segmentation result, and line of sight estimation is performed based on the pupil area; or, at least one of the iris area, pupil area and sclera area in the eye area map is determined according to the image segmentation result, and quality assessment of the eye area map is performed based on at least one of the iris area, pupil area and white of the eye area.
[0156] For example, in a biometric image processing task in an eye area application scenario, the effective image segmentation result can be used for iris recognition to determine the user identity corresponding to the iris. For another example, in a line of sight estimation image processing task in an eye area application scenario, the user's line of sight direction can be estimated in combination with the position information of the pupil area in the iris area. For another example, in an eye image quality assessment in an eye area application scenario, the images of the iris area, pupil area, and sclera area in the eye area map are evaluated separately, and a comprehensive quality assessment of the eye area map is performed based on the evaluation results. This application does not limit the specific application scenarios of the image segmentation results.
[0157] In some embodiments, the detection target is a hand, and the target area map is a hand area map; the palm area in the hand area map is determined according to the image segmentation result, and identity recognition is performed based on the palm area; or, the palm area and finger area in the hand area map are determined according to the image segmentation result, and gesture estimation is performed based on the palm area and finger area.
[0158] In some embodiments, the detection target is the head, and the target area map is the head area map; the facial area in the head area map is determined according to the image segmentation result, and identity recognition is performed based on the facial area; or, the head area and facial area in the head area map are determined according to the image segmentation result, and head posture estimation is performed based on the head area and facial area.
[0159] In some embodiments, the detection target is a road, and the target area map is a road area map; the road object area in the road area map is determined based on image segmentation, and road condition analysis is performed based on the road object area; or, the vehicle area in the road area map is determined based on image segmentation, and traffic control is performed based on the vehicle area.
[0160] Through the above method, according to business needs, using image segmentation results with better segmentation quality for processing helps to improve the quality and efficiency of completing subsequent image processing tasks.
[0161] In this application, the uncertainty estimation of the image segmentation results does not require the use of the segmentation results of another auxiliary segmentation model, but instead performs uncertainty estimation based on the feature information of the target area map, thereby avoiding the problem of poor accuracy of uncertainty estimation caused by inaccurate segmentation results of the auxiliary segmentation model, and improving the accuracy of uncertainty estimation.
[0162] Furthermore, by extracting the target region from the image, a target region map (which can be understood as a sub-image of the image) is obtained. Based on the feature information of the target region map, a confidence assessment is performed on each pixel in the target region map. Then, based on the confidence of each pixel, the uncertainty of the image segmentation result is obtained. Because the target region map is a sub-image of the image, it is obtained by removing most of the non-target areas in the image to include the target region. Therefore, it can remove interference from other image content in the image besides the target region, improve the accuracy of the pixel confidence estimation, and thus improve the accuracy of the uncertainty estimation of the image segmentation result.
[0163] In addition, since the target area map is an image obtained by removing other image contents except the target area in the image, performing image segmentation only on the extracted target area map helps to reduce the amount of calculation, thereby improving the efficiency of image segmentation, and the image segmentation will not be interfered by factors such as the above-mentioned other image contents, thereby improving the accuracy of the image segmentation results.
[0164] The following describes the training process for the image segmentation model and confidence estimation model used in the image processing method. The usage process for the image segmentation model and confidence estimation model corresponds to the method steps in the training process. For details not described in detail in the embodiments corresponding to the training process, please refer to the description of the embodiments corresponding to the usage process. Similarly, for details not described in detail in the embodiments corresponding to the usage process, please refer to the description of the embodiments corresponding to the training process.
[0165] Please refer to Figure 4 , which shows a flow chart of a model training method provided by an embodiment of the present application. The execution subject of each step of the method can be a computer device, for example, the computer device can be Figure 1 The model training device 10 in the solution implementation environment shown. The method may include at least one of the following steps (410-450):
[0166] Step 410 : extracting a target area from the first sample image to obtain a first target area map, wherein pixels in the target area respectively belong to one of a plurality of preset pixel categories.
[0167] In this embodiment, the first sample image is a sample image used to train the confidence estimation model, the second sample image is a sample image used to train the image processing model, and the third sample image is a sample image used to train the object detection model. It should be noted that the first sample image, the second sample image, and the third sample image can be the same or different, and this application does not impose any restrictions on this.
[0168] In some embodiments, the sample images refer to images used for training, which may come from an open source dataset or an artificially created dataset, and this application does not limit this.
[0169] In some embodiments, the sample image has corresponding annotated image segmentation results and annotated position information, the annotated image segmentation results are used to indicate the annotated preset pixel category of each pixel in the sample image, and the annotated position information is used to indicate the annotated position of the target area in the sample image.
[0170] In some embodiments, there may be multiple sample images.
[0171] In some embodiments, before executing step 410, the third sample image is input into the target detection model, and the target detection model outputs the predicted position information of the target area in the third sample image; based on the difference between the predicted position information and the labeled position information of the target area, the loss function value of the target detection model is obtained; based on the loss function value of the target detection model, the parameters of the target detection model are adjusted to obtain a trained target detection model, wherein the trained target detection model is used to extract the target area in the first sample image to obtain a first target area map.
[0172] The target detection model is a neural network model used to identify the target area in the sample image. It can be a model built based on a convolutional neural network, a fully connected neural network, a feedforward neural network, etc., which is not limited in this application.
[0173] In some embodiments, a sample image is input into a target detection model to obtain spatial features of the sample image; the target detection model analyzes and processes the spatial features of the sample image and outputs predicted position information of the target area in the sample image; the loss function value of the target detection model is determined based on the predicted position information and the annotated position information of the target area; and the parameters of the target detection model are adjusted based on the loss function value.
[0174] For example, Figure 5As shown, it shows a flowchart of the target detection model training provided by one embodiment of the present application. The training process of the target detection model is divided into 6 stages: the third sample image preparation stage, the target detection stage, the loss function value calculation stage, the stage of judging whether the training is terminated, the target detection model optimization stage and the training completion stage. The workflow of these 6 stages is as follows:
[0175] (1) Third sample image preparation stage: multiple third sample images are obtained from the dataset to form a batch of training data;
[0176] (2) Target detection stage: The third sample image is input into the target detection model. The target detection model processes the sample image using a variety of computational operations, such as convolution, ReLU, and pooling, to extract key features from the third sample image and output the predicted location information of the target area.
[0177] (3) Loss function value calculation stage: The loss function value of the target detection model can be calculated using loss functions such as classification function (such as softmax), cross entropy loss, mean squared error (MSE), intersection over union (IoU), and other deformation loss functions of IoU. For example, the formula for the loss function value of the target detection model calculated using the intersection over union loss function is as follows:
[0178]
[0179] Among them, L IoU represents the loss function value of the target detection model calculated using the intersection-over-union loss function, the predicted position information indicates the predicted target area in the third sample image, and the annotated position information indicates the annotated target area in the third sample image;
[0180] (4) Determine whether the training is terminated: Determine whether the loss function value calculated in the previous step meets the pre-set training termination condition. The training termination condition can be one of the following conditions: reaching the specified maximum number of iterations, the loss function value is lower than a preset threshold, etc. When the termination condition is met, execute step (6); when the termination condition is not met, execute step (5);
[0181] (5) Target detection model optimization stage: The parameters of the target detection model are adjusted based on the gradient descent method. The gradient descent method can be one of the following methods: stochastic gradient descent, stochastic gradient descent with momentum term, Adam optimizer, and Adagrad optimizer. After adjusting the parameters of the target detection model, the process returns to step (1) to continue execution.
[0182] (6) Completion of the training phase: Terminate the training of the target detection model and obtain the trained target detection model.
[0183] After completing the training of the target detection model, the parameters of the target detection model are fixed and no longer changed, and then other subsequent training steps are continued.
[0184] Through the above method, a target detection model with good recognition performance for the target area is trained, which helps to improve the accuracy and robustness when processing sample images.
[0185] In some embodiments, after completing the training of the target detection model, the target area in the second sample image is extracted to obtain a second target area map; the second target area map is input into the image segmentation model, and the image segmentation model outputs the predicted image segmentation result of the second target area map; based on the difference between the predicted image segmentation result and the labeled image segmentation result of the second target area map, the loss function value of the image segmentation model is obtained; based on the loss function value of the image segmentation model, the parameters of the image segmentation model are adjusted to obtain the trained image segmentation model.
[0186] In some embodiments, the image segmentation model includes a feature extraction network and a feature classification network, the feature extraction network is used to extract feature information, and the feature classification network is used to assign a preset pixel category to each pixel in the target area map.
[0187] In some embodiments, the parameters of the image segmentation model can be adjusted using a gradient descent method, for example, stochastic gradient descent, stochastic gradient descent with momentum, Adam optimizer, Adagrad optimizer, and the like.
[0188] In some embodiments, based on the difference between the predicted image segmentation result corresponding to the second target area map and the labeled image segmentation result of the second target area map, the loss function value of the image segmentation model is calculated; according to the loss function value of the image segmentation model, the parameters of the image segmentation model are adjusted to obtain the trained image segmentation model.
[0189] For example, Figure 6 As shown, it shows a flowchart of the image segmentation model training provided by one embodiment of the present application. The training process of the image segmentation model is divided into 7 stages: the second target area map preparation stage, the feature extraction stage, the feature classification stage, the loss function value calculation stage, the training termination determination stage, the image segmentation model optimization stage and the training completion stage. The workflow of these 7 stages is as follows:
[0190] (1) Second target region map preparation stage: multiple second sample images are obtained from the data set to form a batch of training data, and based on the trained target detection model, a second target region map corresponding to the second sample images is obtained;
[0191] (2) Feature extraction stage: The second target region map is input into the feature extraction network of the image segmentation model. The feature extraction network processes the second target region map using a variety of computational operations, such as convolution, ReLU, and pooling, to extract feature information from the second target region map.
[0192] (3) Feature classification stage: The feature information obtained in the previous step is input into the feature classification network in the image segmentation model. The feature classification network assigns a preset pixel category to each pixel in the second target area image and obtains the corresponding predicted image segmentation result;
[0193] (4) Loss function value calculation stage: The loss function value of the image segmentation model can be calculated using loss functions such as classification function (such as softmax), cross entropy loss, variant cross entropy loss, and classification function that introduces category balance. For example, the formula for calculating the loss function value of the image segmentation model using the weighted cross entropy loss function is as follows:
[0194]
[0195] Among them, L cross represents the loss function value obtained by calculating the image segmentation model using the weighted cross entropy loss function, M represents the number of preset pixel categories, c represents one of the M preset pixel categories, ω c Represents the weight of the preset pixel category c, y c Indicates whether pixel x belongs to the preset pixel category c, y c Based on the segmentation result of the labeled image, x represents any pixel in the second target area map, p c Indicates the predicted probability value of pixel x for the preset pixel category c, p c Obtaining a predicted image segmentation result corresponding to the second target region map;
[0196] (5) Determine whether the training is terminated: Determine whether the loss function value calculated in the previous step meets the pre-set training termination condition. The training termination condition can be one of the following conditions: reaching a specified maximum number of iterations, the loss function value is lower than a preset threshold, etc. When the termination condition is met, execute step (7); when the termination condition is not met, execute step (6);
[0197] (6) Image segmentation model optimization stage: The parameters of the image segmentation model are adjusted based on the gradient descent method. The gradient descent method can be one of the following methods: stochastic gradient descent, stochastic gradient descent with momentum term, Adam optimizer, and Adagrad optimizer. After adjusting the parameters of the image segmentation model, the process returns to step (1) to continue execution.
[0198] (7) Completion of the training phase: Terminate the training of the image segmentation model and obtain the trained image segmentation model.
[0199] After completing the training of the image segmentation model, the parameters of the image segmentation model are fixed and no longer changed, and then other subsequent training steps are continued.
[0200] Through the above method, an image segmentation model with better image segmentation effect is trained. The trained image segmentation model is used for more accurate image segmentation, can adapt to images with different interference factors, and improve the accuracy of the segmentation results of images containing more interference factors.
[0201] In step 420 , the first target region map is input into the feature extraction network of the trained image segmentation model, and the feature extraction network of the trained image segmentation model outputs feature information of the first target region map.
[0202] The process of extracting features from the target area map during training is the same as the process of extracting features from the target area map during use. For details, please refer to Figure 2 This corresponds to the part of step 220 in the embodiment.
[0203] Step 430: Input the feature information into the feature classification network of the trained image segmentation model, and the feature classification network of the trained image segmentation model outputs the image segmentation result of the first target area map, where the image segmentation result indicates the preset pixel category to which each pixel in the first target area map belongs.
[0204] The process of segmenting the target region map to obtain the predicted image segmentation result during training is the same as the process of segmenting the target region map to obtain the target segmentation result during use. For details, please refer to Figure 2 This corresponds to step 230 in the embodiment.
[0205] In step 440, the feature information is input into a confidence estimation model, and the confidence estimation model obtains the confidence of the preset pixel category to which each pixel in the first target area image belongs. The confidence of the preset pixel category to which each pixel in the first target area image belongs is used to obtain the uncertainty of the image segmentation result.
[0206] In some embodiments, feature information is input into a confidence estimation model, and the confidence estimation model outputs confidence information, the confidence information including a confidence feature matrix of each pixel in the first target area map, the confidence feature matrix of the pixel is used to determine the confidence of the preset pixel category to which the pixel belongs; based on the confidence feature matrix of each pixel in the first target area map, the confidence of the preset pixel category to which each pixel in the first target area map belongs is obtained.
[0207] In the above manner, the confidence estimation model is used to perform confidence estimation on each pixel in the first target area map. The accuracy of the confidence information obtained can reflect the accuracy of the confidence estimation of the confidence estimation model. Therefore, the parameters of the confidence estimation model can be adjusted based on the obtained confidence information.
[0208] The process of calculating the uncertainty of the image segmentation results during training corresponds to the process of calculating the uncertainty of the image segmentation results during usage. For details, please refer to Figure 2 This corresponds to the portion of step 250 in the embodiment.
[0209] Step 450 : Adjust the parameters of the confidence estimation model according to the feature information and the confidence of the preset pixel category to which each pixel in the first target area map belongs, to obtain a trained confidence estimation model.
[0210] In some embodiments, feature information and image segmentation results corresponding to the sample image are obtained based on the trained target detection model and the trained image segmentation model.
[0211] In some embodiments, the class center information of each preset pixel category among multiple preset pixel categories is determined, wherein the class center information of each preset pixel category is used to characterize the characteristics of each preset pixel category; based on the characteristic information, the class center information of each preset pixel category, and the confidence of the preset pixel category to which each pixel in the first target area map belongs, the loss function value of the confidence estimation model is obtained; based on the loss function value of the confidence estimation model, the parameters of the confidence estimation model are adjusted to obtain the trained confidence estimation model.
[0212] For each preset pixel category, the class center information of the preset pixel category is used to characterize the characteristics of the preset pixel category, which can be understood as the overall characteristics or average characteristics of the preset pixel category.
[0213] In some embodiments, there are multiple first sample images, and the class center information of each preset pixel category is determined through the image segmentation results and feature information of the multiple first sample images.
[0214] By adjusting the parameters of the confidence estimation model according to the calculated loss function value in the above manner, the performance of the confidence estimation model for pixel confidence estimation can be effectively improved.
[0215] In some embodiments, there are two methods for determining the class center information of each preset pixel class:
[0216] (1) obtaining a feature vector of each pixel in T reference region images, wherein the pixels in the reference region images have multiple preset pixel categories, and the preset pixel categories to which each pixel in the reference region images belongs are known, and T is a positive integer; determining class center information of each preset pixel category in the multiple preset pixel categories based on the feature vector of each pixel in the T reference region images and the preset pixel category to which each pixel in the T reference region images belongs;
[0217] (2) The parameters of the image segmentation model include feature weight vectors of each preset pixel category in a plurality of preset pixel categories, and the feature weight vectors of each preset pixel category are used to perform operations with the feature vectors of pixels in the feature information to obtain a probability value that the pixel belongs to each preset pixel category, and the probability value that the pixel belongs to each preset pixel category is used to determine the image segmentation result; the parameters of the trained image segmentation model are parsed to obtain updated feature weight vectors of each preset pixel category, wherein the updated feature weight vectors of each preset pixel category are obtained by adjusting the initial feature weight vectors of each preset pixel category during the training process of the image segmentation model; based on the updated feature weight vectors of each preset pixel category, the class center information of each preset pixel category is determined.
[0218] In some embodiments, the feature weight vectors of each preset pixel category are respectively inner-producted with the feature vectors of the pixels in the feature information to obtain the inner-product operation results corresponding to each preset pixel category; the inner-product operation results corresponding to each preset pixel category are activated to obtain the probability value of the pixel belonging to each preset pixel category.
[0219] Exemplarily, assuming that the size of the feature vector corresponding to the pixel is 1*D, D is a positive integer, the preset pixel categories corresponding to the pixel are preset pixel category A and preset pixel category B, and the size of the feature weight vector corresponding to the preset pixel category A and the preset pixel category B is also 1*D. The feature weight vectors corresponding to the preset pixel category A and the preset pixel category B are respectively inner-producted with the above feature vectors to obtain two inner-product operation results. The softmax function is used to activate the above two inner-product operation results to obtain the probability value of the pixel belonging to the preset pixel category A and the probability value of the pixel belonging to the preset pixel category B.
[0220] In some embodiments, the T reference region maps are obtained by inputting T first sample images into a target detection model.
[0221] In some embodiments, for method (1), for each preset pixel category, based on the labeled segmentation results of T reference area maps, the feature vectors corresponding to each pixel belonging to the preset pixel category are obtained from the feature information of each of the T sample images, where T is a positive integer; the feature vectors corresponding to each pixel belonging to the preset pixel category are averaged to obtain the class center information of the preset pixel category.
[0222] In some embodiments, for method (2), the parameters of the trained image segmentation model are parsed to obtain feature weight vectors for each preset pixel category, and the feature weight vectors for each preset pixel category are used to classify each pixel in the target area map; for each preset pixel category, the feature weight vector of the preset pixel category is normalized to obtain class center information of the preset pixel category. Wherein, normalization refers to dividing the feature weight vector by the modulus of the feature weight vector.
[0223] Exemplarily, for method (1), assuming that there are T first sample images in total, the size of the reference area map corresponding to each first sample image is 100×200 pixels, and the feature information includes D features, then the size of the feature information of each image is 100×200×D, and assuming that the number of preset pixel categories included in the labeled image segmentation result is C. For any preset pixel category c among the C preset pixel categories, the feature vectors of all pixels belonging to the preset pixel category c in the above T first sample images are taken out, and these taken out feature vectors are averaged to obtain the class center information of the preset pixel category c, which is 1×D in size.
[0224] Exemplarily, for method (2), assuming that there are T first sample images, the size of the target area map corresponding to each first sample image is 100×200 pixels, the number of preset pixel categories included in the image segmentation result is C, the feature information includes D features, and the feature weight vectors of each preset pixel category are obtained by parsing from the trained image segmentation model. The size of the feature weight vector of each preset pixel category is 1×D. For any preset pixel category c among the C preset pixel categories, the feature weight vector of the preset pixel category c is normalized to obtain the class center information of the preset pixel category c, which is 1×D in size. The above-mentioned normalization process refers to dividing the feature weight vector by the modulus of the feature weight vector.
[0225] The above method provides two methods for obtaining class center information, among which method (2) is simpler and more efficient. After the image segmentation model training is completed, the class center information of each preset pixel category can be directly obtained without additional calculation.
[0226] In some embodiments, for each pixel in the first target area map, the feature vector of the pixel in the feature information is subtracted from the class center information of the preset pixel category to which the pixel belongs to obtain a first difference; based on the confidence of the preset pixel category to which the pixel belongs and the first difference, the confidence loss of the pixel is obtained; based on the confidence loss of each pixel in the first target area map, the loss function value of the confidence estimation model is obtained.
[0227] Exemplarily, for any pixel x in the first target region map, the formula for calculating the confidence loss of pixel x is as follows:
[0228] L x =||tr(∑ x )-(f x -w x∈c )·(f x -w x∈c )||
[0229] Among them, L x represents the confidence loss of pixel x, tr(*) represents the trace of matrix *, (a)·(b) represents the inner product of vector a and vector b, ||*|| represents the bi-norm of *, c represents one of the M preset pixel categories, w x∈c Represents the class center information of the preset pixel category c, ∑ x Represents the confidence feature matrix of pixel x, f x The feature vector representing pixel x.
[0230] In some embodiments, the confidence losses corresponding to each pixel in the first target area map can be added together to obtain the loss function value of the confidence estimation model; or the confidence losses corresponding to each pixel in the first target area map can be averaged to obtain the loss function value of the confidence estimation model.
[0231] Through the above method, based on the class center information of a preset pixel class, the distance between the pixel's feature vector and the preset pixel class can be determined, thereby accurately assessing the confidence loss of the pixel. This confidence loss reflects the accuracy of the confidence assessment for the preset pixel class to which the pixel belongs. In this way, based on the accurate confidence loss, the accuracy of the loss function value calculated for the confidence estimation model can be ensured, thereby improving the training effect of the model.
[0232] For example, Figure 7As shown, it shows a flowchart of the confidence estimation model training provided by one embodiment of the present application. The training process of the confidence estimation model is divided into 7 stages: feature information preparation stage, class center information acquisition stage, uncertainty estimation stage, loss function value calculation stage, judgment of whether training is terminated stage, confidence estimation model optimization stage and completion training stage. The workflow of these 7 stages is as follows:
[0233] (1) Feature information preparation stage: multiple first sample images are obtained from the dataset to form a batch of training data, and feature information corresponding to the first sample images is obtained based on the trained object detection model and the trained image segmentation model;
[0234] (2) Class center information acquisition stage: Based on one of the two methods for obtaining class center information, the class center information of each preset pixel category of the image segmentation result is obtained;
[0235] (3) Confidence estimation stage: The obtained feature information is input into the confidence estimation model, and the confidence estimation model outputs confidence information;
[0236] (4) Loss function value calculation stage: The loss function value of the image segmentation model is calculated based on feature information, confidence information and class center information. The calculation formula of the loss function value is as follows:
[0237]
[0238] Among them, L uncertainty Represents the loss function value of the confidence estimation model, L x represents the confidence loss of pixel x;
[0239] (5) Determine whether the training is terminated: Determine whether the loss function value calculated in the previous step meets the pre-set training termination condition. The training termination condition can be one of the following conditions: reaching a specified maximum number of iterations, the loss function value is lower than a preset threshold, etc. When the termination condition is met, execute step (7); when the termination condition is not met, execute step (6);
[0240] (6) Confidence estimation model optimization stage: The parameters of the confidence estimation model are adjusted based on the gradient descent method. The gradient descent method can be one of the following methods: stochastic gradient descent, stochastic gradient descent with momentum term, Adam optimizer, and Adagrad optimizer. After adjusting the parameters of the confidence estimation model, the process returns to step (1) to continue execution.
[0241] (7) Completion of the training phase: terminating the training of the confidence estimation model to obtain the trained confidence estimation model.
[0242] To summarize, the technical solution provided by the embodiment of the present application obtains the first target area map from the first sample image by extracting the first target area map, inputs the first target area map into the feature extraction network of the trained image segmentation model, obtains the feature information of the first target area map, inputs the feature information into the feature classification network of the trained image segmentation model, obtains the image segmentation result corresponding to the first target area map, and then obtains the trained image segmentation model, obtains the image segmentation result and feature information based on the trained image segmentation model, inputs the feature information into the confidence estimation model, obtains the uncertainty corresponding to the image segmentation result, adjusts the parameters of the confidence estimation model according to the feature information and the confidence of each pixel in the first target area map, and obtains the trained confidence estimation model, thereby improving the accuracy and robustness of image segmentation of the image segmentation model and improving the accuracy of the uncertainty estimation of the image segmentation result by the confidence estimation model by sequentially training the image segmentation model and the confidence estimation model.
[0243] Based on the understanding of the above embodiments, the following uses a specific embodiment applied in a biometric application scenario to illustrate the use process of the image processing method and model training method of the present application.
[0244] For example, Figure 8 As shown, it shows a flowchart of the use of the image processing method and model training method provided by an embodiment of the present application, which is divided into two stages: the model training stage and the model use stage. In the model training stage, the above-mentioned model training method is used to train the eye detection model, iris segmentation model and confidence estimation model in sequence to obtain the trained eye detection model, the trained iris segmentation model and the trained confidence estimation model. In the model use stage, the above-mentioned iris segmentation method is used, combined with the trained eye detection model, the eye area clipping part, the trained iris segmentation model, the trained confidence estimation model and the threshold comparison part, to output the iris segmentation result. Among them, the eye detection model, which is a specific implementation of the target detection model introduced above, is used to obtain the position information of the eye area from the image. The iris segmentation model, which is a specific implementation of the image segmentation model introduced above, is used to segment the iris area from the eye area map. The eye area clipping part is used to clip the eye area map from the image based on the position information of the eye area.
[0245] Based on the previous example, the application process of each model in the model usage phase is as follows Figure 9As shown, it shows a flowchart of the use of the model provided by an embodiment of the present application after deployment, and the image is input into the eye detection model. Then the position information of the eye area is input into the eye area clipping part to obtain the corresponding eye area map. Next, the eye area map is input into the feature extraction network of the iris segmentation model to obtain feature information. By inputting the feature information into the feature classification network and the confidence estimation model of the iris segmentation model respectively, the iris segmentation result and the uncertainty of the iris segmentation result are obtained respectively. In the threshold comparison part, based on the uncertainty of the iris segmentation result, it is determined whether the obtained iris segmentation result is valid. In the biometric image processing task of the eye area application scenario, if it is valid, the iris area in the iris segmentation result is subjected to feature extraction, comparative analysis and other processing to determine the identity of the user; in the line of sight estimation image processing task of the eye area application scenario, the corresponding pupil area is determined from the iris segmentation result, and the pupil area is analyzed in combination with other features to obtain the corresponding line of sight estimation result.
[0246] For the hand area application scenario, the usage process of the image processing method and model training method is basically the same as the above example. After obtaining the image segmentation result, if the image segmentation result is valid, then in the biometric image processing task in the hand area application scenario, the corresponding palm area is determined from the image segmentation result, and the palm area is subjected to feature extraction, comparative analysis and other processing to determine the identity of the user; in the hand posture recognition image processing task of the hand area application scenario, image processing is performed on the image containing the hand area, and an image containing each finger area and palm area is extracted from the image. The shape, size and position of the finger area contained in the above image are analyzed to estimate the corresponding hand posture.
[0247] For the head region application scenario, the image processing method and model training method are basically the same as the above example. After obtaining the image segmentation result, if the image segmentation result is valid, the facial region is determined from the image segmentation result in the biometric image processing task of the head region application scenario. The facial recognition model can be used to perform facial recognition on the facial region to determine the user's identity. The facial recognition model can be any AI model used for facial recognition.
[0248] For the road area application scenario, the image processing and model training methods are essentially the same as in the previous example. After obtaining the image segmentation results, if they are valid, the traffic control image processing task for the road area application scenario uses them to determine the vehicle area. A traffic control algorithm is then used to analyze the current road conditions and control traffic flow. The traffic control algorithm can be any algorithm for analyzing and controlling vehicle traffic flow.
[0249] This application can also be applied to other scenarios. For specific implementation methods, please refer to the above embodiments and will not be described in detail here.
[0250] The following are device embodiments of the present application, which can be used to implement the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.
[0251] Please refer to Figure 10 , which shows a block diagram of an image processing device provided by an embodiment of the present application. The device has the function of implementing the above-mentioned image processing method, and the function can be implemented by hardware or by hardware executing corresponding software. The device can be the model using device 20 introduced above, or it can be set in the model using device 20. Figure 10 As shown, the apparatus 1000 may include a region extraction module 1010 , a feature extraction module 1020 , an image segmentation module 1030 , a confidence determination module 1040 and an uncertainty estimation module 1050 .
[0252] The region extraction module 1010 is configured to extract a target region in an image and obtain a target region map, wherein pixels in the target region respectively belong to one of a plurality of preset pixel categories.
[0253] The feature extraction module 1020 is configured to extract features from the target area map to obtain feature information of the target area map.
[0254] The image segmentation module 1030 is configured to perform image segmentation on the target area map based on the feature information to obtain an image segmentation result of the target area map, wherein the image segmentation result indicates the preset pixel category to which each pixel in the target area map belongs.
[0255] The confidence determination module 1040 is configured to obtain the confidence of the preset pixel category to which each pixel in the target area map belongs based on the feature information.
[0256] The uncertainty estimation module 1050 is configured to obtain the uncertainty of the image segmentation result according to the confidence of the preset pixel category to which each pixel in the target area map belongs.
[0257] In some embodiments, the confidence determination module 1040 includes a confidence information acquisition submodule and a confidence acquisition submodule (in Figure 10 not shown).
[0258] A confidence information acquisition submodule is used to input the feature information into a confidence estimation model, and the confidence estimation model outputs confidence information. The confidence information includes a confidence feature matrix of each pixel in the target area map, and the confidence feature matrix of the pixel is used to determine the confidence of the preset pixel category to which the pixel belongs.
[0259] The confidence acquisition submodule is used to obtain the confidence of the preset pixel category to which each pixel in the target area map belongs according to the confidence feature matrix of each pixel in the target area map.
[0260] In some embodiments, the confidence acquisition submodule is used to determine, for each pixel in the target area map, the trace of the confidence feature matrix of the pixel as the confidence of the preset pixel category to which the pixel belongs.
[0261] In some embodiments, the uncertainty estimation module 1050 is configured to obtain the uncertainty of the image segmentation result based on the sum of the logarithms of the confidences of the preset pixel categories to which each pixel in the target area map belongs.
[0262] In some embodiments, the apparatus 1000 further includes a post-processing module (in Figure 10 (not shown) is used to perform subsequent image processing tasks according to the image segmentation result when the uncertainty is less than a preset threshold.
[0263] In some embodiments, the target area map is an eye area map of the target object; the post-processing module is used to determine the iris area in the eye area map according to the image segmentation result, and identify the identity of the target object based on the iris area; or, determine the pupil area in the eye area map according to the image segmentation result, and estimate the line of sight direction of the target object based on the pupil area; or, determine at least one of the iris area, pupil area and white of the eye area in the eye area map according to the image segmentation result, and estimate the image quality of the eye area map based on at least one of the iris area, pupil area and white of the eye area.
[0264] In some embodiments, the region extraction module 1010 is used to input the image into a target detection model, and the target detection model outputs the location information of the target region in the image; based on the location information, the target region is cut out from the image to obtain the target region map.
[0265] In this application, the uncertainty estimation of the image segmentation results does not require the use of the segmentation results of another auxiliary segmentation model, but instead performs uncertainty estimation based on the feature information of the target area map, thereby avoiding the problem of poor accuracy of uncertainty estimation caused by inaccurate segmentation results of the auxiliary segmentation model, and improving the accuracy of uncertainty estimation.
[0266] Please refer to Figure 11 , which shows a block diagram of a model training device provided by an embodiment of the present application. The device has the function of implementing the above-mentioned model training method, and the function can be implemented by hardware or by hardware executing corresponding software. The device can be the model training device 10 introduced above, or it can be set in the model training device 10. Figure 11 As shown, the apparatus 1100 may include a sample target region extraction module 1110 , a feature information extraction module 1120 , an image segmentation processing module 1130 , a confidence information acquisition module 1140 and an estimation model adjustment module 1150 .
[0267] The target region extraction module 1110 is configured to extract a target region from the first sample image to obtain a first target region map, wherein pixels in the target region respectively belong to one of a plurality of preset pixel categories.
[0268] The feature information extraction module 1120 is used to input the first target area map into the feature extraction network of the trained image segmentation model, and the feature extraction network of the trained image segmentation model outputs the feature information of the first target area map.
[0269] The image segmentation processing module 1130 is used to input the feature information into the feature classification network of the trained image segmentation model, and the feature classification network of the trained image segmentation model outputs the image segmentation result of the first target area map, and the image segmentation result indicates the preset pixel category to which each pixel in the first target area map belongs.
[0270] The confidence information acquisition module 1140 is used to input the feature information into a confidence estimation model, and obtain the confidence of the preset pixel category to which each pixel in the first target area map belongs from the confidence estimation model. The confidence of the preset pixel category to which each pixel in the first target area map belongs is used to obtain the uncertainty of the image segmentation result.
[0271] The estimation model adjustment module 1150 is used to adjust the parameters of the confidence estimation model according to the feature information and the confidence of the preset pixel category to which each pixel in the first target area map belongs, to obtain a trained confidence estimation model.
[0272] In some embodiments, the estimation model adjustment module 1150 includes a class center determination submodule, a loss calculation submodule, and a model adjustment submodule (in Figure 11 not shown).
[0273] The class center determination submodule is configured to determine class center information of each preset pixel class among the plurality of preset pixel classes, wherein the class center information of each preset pixel class is used to characterize features of each preset pixel class.
[0274] The loss calculation submodule is used to obtain the loss function value of the confidence estimation model based on the feature information, the class center information of each preset pixel category, and the confidence of the preset pixel category to which each pixel in the first target area map belongs.
[0275] The model adjustment submodule is used to adjust the parameters of the confidence estimation model according to the loss function value of the confidence estimation model to obtain the trained confidence estimation model.
[0276] In some embodiments, the class center determination submodule is used to obtain the feature vectors of each pixel in T reference area maps, where the pixels of the reference area maps have the multiple preset pixel categories, and the preset pixel categories to which each pixel in the reference area maps belongs are known, and T is a positive integer; based on the feature vectors of each pixel in the T reference area maps and the preset pixel categories to which each pixel in the T reference area maps belongs, the class center information of each preset pixel category in the multiple preset pixel categories is determined.
[0277] In some embodiments, the parameters of the image segmentation model include feature weight vectors of each preset pixel category in the multiple preset pixel categories, and the feature weight vectors of each preset pixel category are used to perform operations with the feature vectors of the pixels in the feature information to obtain probability values that the pixels belong to the respective preset pixel categories, and the probability values that the pixels belong to the respective preset pixel categories are used to determine the image segmentation results; the class center determination submodule is used to parse the parameters of the trained image segmentation model to obtain updated feature weight vectors of the respective preset pixel categories, wherein the updated feature weight vectors of the respective preset pixel categories are obtained by adjusting the initial feature weight vectors of the respective preset pixel categories during the training process of the image segmentation model; based on the updated feature weight vectors of the respective preset pixel categories, the class center information of the respective preset pixel categories is determined.
[0278] In some embodiments, the loss calculation submodule is used to, for each pixel in the first target area map, subtract the class center information of the preset pixel category to which the pixel belongs from the feature vector of the pixel in the feature information to obtain a first difference; obtain the confidence loss of the pixel based on the confidence of the preset pixel category to which the pixel belongs and the first difference; and obtain the loss function value of the confidence estimation model based on the confidence loss of each pixel in the first target area map.
[0279] In some embodiments, the confidence information acquisition module 1140 is used to input the feature information into a confidence estimation model, and the confidence estimation model outputs confidence information, wherein the confidence information includes a confidence feature matrix of each pixel in the first target area map, and the confidence feature matrix of the pixel is used to determine the confidence of the preset pixel category to which the pixel belongs; based on the confidence feature matrix of each pixel in the first target area map, the confidence of the preset pixel category to which each pixel in the first target area map belongs is obtained.
[0280] In some embodiments, the apparatus 1100 further includes a segmentation model training module (in Figure 11 ), and is used to: extract the target area in the second sample image to obtain a second target area map; input the second target area map into an image segmentation model, and output a predicted image segmentation result of the second target area map from the image segmentation model; obtain a loss function value of the image segmentation model based on the difference between the predicted image segmentation result and the labeled image segmentation result of the second target area map; and adjust the parameters of the image segmentation model based on the loss function value of the image segmentation model to obtain the trained image segmentation model.
[0281] In some embodiments, the apparatus 1100 further includes a detection model training module (in Figure 11 ), which is not shown in the figure, and is used to: input a third sample image into a target detection model, and the target detection model outputs the predicted position information of the target area in the third sample image; obtain the loss function value of the target detection model based on the difference between the predicted position information and the labeled position information of the target area; adjust the parameters of the target detection model based on the loss function value of the target detection model to obtain a trained target detection model, wherein the trained target detection model is used to extract the target area in the first sample image to obtain the first target area map.
[0282] To summarize, the technical solution provided by the embodiment of the present application obtains the first target area map from the first sample image by extracting the first target area map, inputs the first target area map into the feature extraction network of the trained image segmentation model, obtains the feature information of the first target area map, inputs the target area map into the feature classification network of the trained image segmentation model, obtains the image segmentation result corresponding to the first target area map, obtains the trained image segmentation model, obtains the image segmentation result and feature information based on the trained image segmentation model, inputs the feature information into the confidence estimation model, obtains the uncertainty corresponding to the image segmentation result, adjusts the parameters of the confidence estimation model according to the feature information and the confidence of each pixel in the first target area map, and obtains the trained confidence estimation model, thereby improving the accuracy and robustness of image segmentation of the image segmentation model and improving the accuracy of the uncertainty estimation of the image segmentation result by the confidence estimation model by sequentially training the image segmentation model and the confidence estimation model.
[0283] It should be noted that the apparatus provided in the above embodiments, when implementing its functions, is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0284] Please refer to Figure 12 , which shows a block diagram of a computer device 1200 provided in one embodiment of the present application. The computer device 1200 may be Figure 1 The model training device 10 in the implementation environment shown can also be Figure 1 The model using device 20 in the illustrated implementation environment is used to implement the image processing method or model training method provided in the above embodiments. Specifically:
[0285] Typically, the computer device 1200 includes a processor 1210 and a memory 1220 .
[0286] The processor 1210 may include one or more processing cores, such as a 4-core processor, an 8-core processor, and the like. The processor 1210 may be implemented in at least one hardware form of digital signal processing (DSP), field programmable gate array (FPGA), and programmable logic array (PLA). The processor 1210 may also include a main processor and a coprocessor. The main processor is a processor for processing data in an awake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor for processing data in a standby state. In some embodiments, the processor 1210 may be integrated with a graphics processing unit (GPU), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1210 may also include an AI processor for processing computing operations related to machine learning.
[0287] The memory 1220 may include one or more computer-readable storage media, which may be non-transitory. The memory 1220 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 1220 is used to store a computer program, which is configured to be executed by one or more processors to implement the above-mentioned image processing method or model training method.
[0288] Those skilled in the art will understand that Figure 12 The structure shown in the figure does not constitute a limitation on the computer device 1200, and the computer device 1200 may include more or fewer components than shown in the figure, or combine some components, or adopt a different component arrangement.
[0289] In an exemplary embodiment, a computer-readable storage medium is further provided, wherein a computer program is stored in the storage medium, and the computer program is executed by a processor to implement the above-mentioned image processing method or model training method. Optionally, the computer-readable storage medium may include: a read-only memory (ROM), a random access memory (RAM), a solid-state drive (SSD) or an optical disc, etc. Among them, the random access memory may include a resistance random access memory (ReRAM) and a dynamic random access memory (DRAM).
[0290] In an exemplary embodiment, a computer program product is also provided, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the above-described image processing method or model training method.
[0291] It should be noted that the collection and processing of relevant data (such as pictures, videos, etc.) in this application should be strictly in accordance with the requirements of relevant national laws and regulations when applied in practice, and the informed consent or separate consent of the personal information subject should be obtained. Subsequent data use and processing should be carried out within the scope of authorization of laws and regulations and the personal information subject.
[0292] It should be understood that the "multiple" mentioned in this article refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent three situations: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the previous and subsequent associated objects are in an "or" relationship. In addition, the step numbers described in this article only illustrate a possible execution sequence between the steps. In some other embodiments, the above steps may not be executed in the order of the numbers, such as two steps with different numbers are executed at the same time, or two steps with different numbers are executed in the opposite order to the diagram. The embodiments of the present application do not limit this.
[0293] The above description is merely an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. An image processing method, characterized in that: The method comprises: Extracting a target area in the image to obtain a target area map, wherein pixels of the target area respectively belong to one of a plurality of preset pixel categories; Performing feature extraction on the target area map to obtain feature information of the target area map; Performing image segmentation on the target area map based on the feature information to obtain an image segmentation result of the target area map, wherein the image segmentation result indicates the preset pixel category to which each pixel in the target area map belongs; Obtaining, based on the feature information, a confidence level of the preset pixel category to which each pixel in the target area map belongs; The uncertainty of the image segmentation result is obtained according to the confidence of the preset pixel category to which each pixel in the target area map belongs.
2. The method according to claim 1, characterized in that The obtaining, based on the feature information, the confidence level of the preset pixel category to which each pixel in the target area map belongs includes: Inputting the feature information into a confidence estimation model, and having the confidence estimation model output confidence information, wherein the confidence information includes a confidence feature matrix of each pixel in the target area map, and the confidence feature matrix of the pixel is used to determine the confidence of the preset pixel category to which the pixel belongs; The confidence level of the preset pixel category to which each pixel in the target area map belongs is obtained according to the confidence feature matrix of each pixel in the target area map.
3. The method according to claim 2, characterized in that The obtaining, based on the confidence feature matrix of each pixel in the target area map, the confidence level of the preset pixel category to which each pixel in the target area map belongs, includes: For each pixel in the target area map, the trace of the confidence feature matrix of the pixel is determined as the confidence of the preset pixel category to which the pixel belongs.
4. The method according to any one of claims 1 to 3, characterized in that Obtaining the uncertainty of the image segmentation result according to the confidence of the preset pixel category to which each pixel in the target area map belongs includes: The uncertainty of the image segmentation result is obtained according to the sum of the logarithms of the confidences of the preset pixel categories to which each pixel in the target area image belongs.
5. The method according to any one of claims 1 to 4, characterized in that After obtaining the uncertainty of the image segmentation result based on the confidence of the preset pixel category to which each pixel in the target area map belongs, the method further includes: When the uncertainty is less than a preset threshold, subsequent image processing tasks are performed according to the image segmentation result.
6. The method according to claim 5, characterized in that The target area map is an eye area map of the target object; The performing subsequent image processing tasks according to the image segmentation result includes: determining an iris region in the eye region map according to the image segmentation result, and identifying the identity of the target object based on the iris region; or, determining a pupil area in the eye area map according to the image segmentation result, and estimating a sight direction of the target object based on the pupil area; or, At least one of the iris area, pupil area and white eye area in the eye area map is determined according to the image segmentation result, and the image quality of the eye area map is estimated based on at least one of the iris area, pupil area and white eye area.
7. The method according to any one of claims 1 to 6, characterized in that The step of extracting the target area from the image to obtain a target area map includes: Inputting the image into a target detection model, and having the target detection model output location information of the target area in the image; Based on the position information, the target area is captured from the image to obtain the target area map.
8. A model training method, characterized in that: The method comprises: Extracting a target area in the first sample image to obtain a first target area map, wherein pixels of the target area respectively belong to one of a plurality of preset pixel categories; Inputting the first target area map into a feature extraction network of a trained image segmentation model, and having the feature extraction network of the trained image segmentation model output feature information of the first target area map; Inputting the feature information into a feature classification network of the trained image segmentation model, and having the feature classification network of the trained image segmentation model output an image segmentation result of the first target area image, wherein the image segmentation result indicates the preset pixel category to which each pixel in the first target area image belongs; Inputting the feature information into a confidence estimation model, obtaining the confidence of the preset pixel category to which each pixel in the first target area image belongs from the confidence estimation model, and using the confidence of the preset pixel category to which each pixel in the first target area image belongs to obtain the uncertainty of the image segmentation result; According to the feature information and the confidence of the preset pixel category to which each pixel in the first target area map belongs, the parameters of the confidence estimation model are adjusted to obtain a trained confidence estimation model.
9. The method according to claim 8, characterized in that The step of adjusting the parameters of the confidence estimation model according to the feature information and the confidence of the preset pixel category to which each pixel in the first target area map belongs to obtain a trained confidence estimation model includes: Determining class center information of each preset pixel class among the plurality of preset pixel classes, wherein the class center information of each preset pixel class is used to characterize features of each preset pixel class; Obtaining a loss function value of the confidence estimation model based on the feature information, the class center information of each preset pixel category, and the confidence of the preset pixel category to which each pixel in the first target area image belongs; According to the loss function value of the confidence estimation model, the parameters of the confidence estimation model are adjusted to obtain the trained confidence estimation model.
10. The method according to claim 9, characterized in that The determining of the class center information of each preset pixel category in the plurality of preset pixel categories includes: Obtaining a feature vector of each pixel in T reference area images, where the pixels of the reference area images have the plurality of preset pixel categories, and the preset pixel category to which each pixel in the reference area images belongs is known, and T is a positive integer; Based on the feature vector of each pixel in the T reference area maps and the preset pixel category to which each pixel in the T reference area maps belongs, class center information of each preset pixel category in the multiple preset pixel categories is determined.
11. The method according to claim 9, characterized in that The parameters of the image segmentation model include a feature weight vector of each preset pixel category in the plurality of preset pixel categories, the feature weight vector of each preset pixel category being used to perform an operation on the feature vector of the pixel in the feature information to obtain a probability value that the pixel belongs to each preset pixel category, and the probability value of the pixel belonging to each preset pixel category being used to determine the image segmentation result; The determining of the class center information of each preset pixel category in the plurality of preset pixel categories includes: Parsing the parameters of the trained image segmentation model to obtain updated feature weight vectors for each of the preset pixel categories, wherein the updated feature weight vectors for each of the preset pixel categories are obtained by adjusting the initial feature weight vectors for each of the preset pixel categories during the training of the image segmentation model; Based on the updated feature weight vectors of the respective preset pixel categories, class center information of the respective preset pixel categories is determined.
12. The method according to any one of claims 9 to 11, characterized in that Obtaining a loss function value of the confidence estimation model based on the feature information, the class center information of each preset pixel category, and the confidence of the preset pixel category to which each pixel in the first target area image belongs, includes: For each pixel in the first target area map, subtracting the class center information of the preset pixel class to which the pixel belongs from the feature vector of the pixel in the feature information to obtain a first difference; Obtaining a confidence loss of the pixel based on the confidence of the preset pixel category to which the pixel belongs and the first difference; Based on the confidence loss of each pixel in the first target area map, a loss function value of the confidence estimation model is obtained.
13. The method according to any one of claims 8 to 12, characterized in that Inputting the feature information into a confidence estimation model, and obtaining, by the confidence estimation model, the confidence of the preset pixel category to which each pixel in the first target area map belongs, includes: Inputting the feature information into a confidence estimation model, and having the confidence estimation model output confidence information, wherein the confidence information includes a confidence feature matrix of each pixel in the first target area map, and the confidence feature matrix of each pixel is used to determine the confidence of the preset pixel category to which the pixel belongs; The confidence level of the preset pixel category to which each pixel in the first target area map belongs is obtained according to the confidence feature matrix of each pixel in the first target area map.
14. The method according to any one of claims 8 to 13, characterized in that The method further comprises: extracting the target area in the second sample image to obtain a second target area map; inputting the second target area map into an image segmentation model, and having the image segmentation model output a predicted image segmentation result of the second target area map; Obtaining a loss function value of the image segmentation model according to a difference between the predicted image segmentation result and the labeled image segmentation result of the second target area map; According to the loss function value of the image segmentation model, the parameters of the image segmentation model are adjusted to obtain the trained image segmentation model.
15. The method according to any one of claims 8 to 14, characterized in that The method further comprises: Inputting the third sample image into a target detection model, and having the target detection model output predicted position information of the target area in the third sample image; Obtaining a loss function value of the target detection model based on a difference between the predicted position information and the annotated position information of the target area; According to the loss function value of the target detection model, the parameters of the target detection model are adjusted to obtain a trained target detection model, wherein the trained target detection model is used to extract the target area in the first sample image to obtain the first target area map.
16. An image processing device, characterized in that: The device comprises: A region extraction module is used to extract a target region in an image and obtain a target region map, wherein pixels of the target region respectively belong to one of a plurality of preset pixel categories; A feature extraction module is used to extract features from the target area map to obtain feature information of the target area map; an image segmentation module, configured to perform image segmentation on the target area map based on the feature information to obtain an image segmentation result of the target area map, wherein the image segmentation result indicates the preset pixel category to which each pixel in the target area map belongs; A confidence determination module, configured to obtain the confidence of the preset pixel category to which each pixel in the target area map belongs based on the feature information; The uncertainty estimation module is used to obtain the uncertainty of the image segmentation result according to the confidence of the preset pixel category to which each pixel in the target area map belongs.
17. A model training device, characterized in that: The device comprises: a target region extraction module, configured to extract a target region from the first sample image to obtain a first target region map, wherein pixels of the target region respectively belong to one of a plurality of preset pixel categories; A feature information extraction module, configured to input the first target area map into a feature extraction network of a trained image segmentation model, and output feature information of the first target area map from the feature extraction network of the trained image segmentation model; an image segmentation processing module, configured to input the feature information into a feature classification network of the trained image segmentation model, and output an image segmentation result of the first target area image from the feature classification network of the trained image segmentation model, wherein the image segmentation result indicates the preset pixel category to which each pixel in the first target area image belongs; a confidence information acquisition module, configured to input the feature information into a confidence estimation model, and obtain, from the confidence estimation model, the confidence of the preset pixel category to which each pixel in the first target area image belongs, wherein the confidence of the preset pixel category to which each pixel in the first target area image belongs is used to obtain the uncertainty of the image segmentation result; The estimation model adjustment module is used to adjust the parameters of the confidence estimation model according to the feature information and the confidence of the preset pixel category to which each pixel in the first target area map belongs, so as to obtain a trained confidence estimation model.
18. A computer device, characterized in that: The computer device includes a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the method according to any one of claims 1 to 7, or to implement the method according to any one of claims 8 to 15.
19. A computer-readable storage medium, characterized in that The storage medium stores a computer program, which is configured to be executed by a processor to implement the method according to any one of claims 1 to 7, or to implement the method according to any one of claims 8 to 15.
20. A computer program product, characterized in that The computer program product comprises a computer program, which is loaded and executed by a processor to implement the method according to any one of claims 1 to 7 or the method according to any one of claims 8 to 15.