Shielding judgment method, electronic equipment, storage medium and program product
By performing two segmentation processing and object repair on the image, the problem of long processing time and inaccurate occlusion judgment caused by the large number of object categories in the image is solved, and more efficient and accurate occlusion judgment is achieved.
Patent Information
- Application Number
- CN202411493164.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-23
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2044-10-23
AI Technical Summary
In the prior art, due to too many object categories contained in the image, the processing time is too long when labeling all objects in the image, which reduces image processing efficiency and is not accurate enough in the occlusion judgment.
By performing two different segmentation processes on the image, the region intersection situation is obtained, and combined with the object repair process, it is determined whether the second object occludes the first object, and the accuracy of occlusion judgment is improved.
More efficient occlusion judgment is achieved, the accuracy and processing efficiency of occlusion judgment are improved, and the occlusion situation of the second object on the first object can be accurately determined.
Smart Images

Figure CN120431323A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular, to an occlusion determination method, an electronic device, a storage medium, and a program product. Background Art
[0002] With the development of image processing technology, after obtaining an image, image processing can be performed on a specified object in the image. For example, the object state (open-eye state or closed-eye state) of the specified object can be recognized. If the specified object in the image is occluded by other objects, occlusion determination can be performed on the other objects before performing image processing on the specified object.
[0003] In the related art, various visual features (such as texture, color, shape, etc.) in the image are extracted to label all objects in the image, and the labeling results of all objects are obtained, so as to determine the occlusion situation of the specified object according to the labeling results.
[0004] However, due to the excessive number of object categories included in the image, the method of labeling all objects in the image will result in too long image processing time and reduce the image processing efficiency. Summary of the Invention
[0005] This application provides an occlusion determination method, an electronic device, and a storage medium, which can perform two different image segmentation processes on a first image, obtain different image segmentation results, and then determine whether a second object in the first image occludes a first object according to the region intersection situation, thereby improving the accuracy of occlusion determination.
[0006] In a first aspect, an occlusion determination method is provided. The method includes: obtaining a first image, where the first image includes a first object and a second object, the first object is in a first region, and the second object is in a second region; performing a first image segmentation process on the first image to obtain a first segmented image, where the first segmented image includes a third region corresponding to the first image, the third region includes the first region, and in the case where the first object is occluded by the second object, a fourth region obtained by performing object repair processing on the first object; determining whether the second object occludes the first object according to the region intersection situation between the second region and the third region, where the region intersection situation is used to indicate an overlapping region formed after the bounding box of the second region overlaps with the bounding box of the third region.
[0007] In the technical solution of this application, after obtaining a first image including a first object and a second object, when the first object and the second object are respectively in a first area and a second area, a first segmentation process is performed on the first image to obtain a first segmented image. Among them, a third area corresponding to the first image is shown in the first segmented image. The third area not only includes the first area, but also includes a fourth area after object repair processing in the case where the first object is occluded. Whether the second object occludes the first object is determined according to the area intersection between the second area and the third area. That is, on the basis of segmenting the first image into foreground and background images, a first segmentation process is also performed on the first image. The object repair function is automatically executed on the occluded first object in the first segmented image, so that the third area corresponding to the first object can be completely reflected in the first segmented image. Thus, according to the area intersection between the second area including only the second object and the first area including only the first object, it is finally determined whether the second object occludes the first object, which can improve the accuracy of the occlusion judgment result.
[0008] It should be understood that the first object and the second object in the first image are only used to distinguish different objects. According to different selected objects, their corresponding occlusion situations are also different. For example: the first object is object a, and the second object is a fan. If object a is used as a reference, then when the fan occludes the face of object a, it can be considered that the second object occludes the first object. If the fan is used as a reference, the fan occludes the face of object a, but the face of object a does not occlude the fan. However, since the hand of object a holds the handle of the fan, the hand of object a occludes the fan.
[0009] In this embodiment, the first object is used as a reference to determine whether the second object occludes the first object.
[0010] Schematically, in the first image, the first object is in the first area, and the second object is in the second area. Among them, the first area and the second area are areas without intersection in the first image.
[0011] Optionally, the first object is completely in the second area, or part of the first object is in the first area and part is in the second area (in this case, part of the area is occluded by the second object). In this embodiment, the first area and the second area are determined only according to the display situation of the first object and the second object in the first image. That is to say, if there is an area a in the first object that is occluded by the second object, then area a will belong to the second area.
[0012] Schematically, the third region includes the first region and the fourth region. As mentioned above, if there is a region a in the first object that is blocked by the second object, then during the first image segmentation process of the first image, it is recognized that region a is blocked by the second object. Ignoring the occlusion of region a by the second object, the other regions of the first object except region a are taken as the first region, region a is taken as the fourth region, and the first region and the fourth region are taken as the third region.
[0013] It should be understood that the third region is used to indicate the complete region of the first object in the first image when there is no occlusion in the first object. The first region may be the same as the third region, or the first region may be smaller than the third region (that is, at this time, there is a partial region of the first object blocked by the second object).
[0014] Schematically, the region intersection situation refers to the overlap situation between the bounding box corresponding to the second region and the bounding box corresponding to the third region. If there is an overlapping region between the bounding box of the second region and the bounding box of the third region, it indicates that the second object occludes the first object. If there is no overlapping region between the bounding box of the second region and the bounding box of the third region, it indicates that the second object does not occlude the first object.
[0015] Combined with the first aspect, in some implementation manners of the first aspect, the above method further includes: obtaining a reference detection box of the first object in the first image, where the reference detection box is used to indicate the position of the first object in the first image; obtaining a first detection box of the second object in the first image according to the region intersection situation between the second region and the third region, where the first detection box is used to indicate the position of the second object in the first image; determining whether the second object occludes the first object according to the overlapping situation between the first detection box and the reference detection box. By the above method, after setting the reference detection box corresponding to the first object in the first image and obtaining the first detection box of the second object in the first image according to the region overlap between the second region and the third region, the edge overlap degree between the reference detection box and the first detection box is calculated, thereby improving the accuracy of occlusion judgment.
[0016] Combined with the first aspect, in some implementation manners of the first aspect, the above method further includes: when there is an overlapping region between the first detection box and the reference detection box, determining the occlusion area of the second object on the first object according to the area of the overlapping region. By the above method, when it is determined that the second object occludes the first object, the occlusion area of the second object on the first object can also be determined according to the area of the overlapping region between the first detection box and the reference detection box. In addition to judging whether there is occlusion, the occlusion area can also be judged, further improving the accuracy of the occlusion judgment result.
[0017] In combination with the first aspect, in some implementations of the first aspect, the above method further includes: performing a second image segmentation process on the first image to obtain a second segmented image, where the second segmented image includes a first region and a second region; performing a morphological erosion process on the second segmented image to obtain an eroded image, where the eroded image includes a fifth region, and the fifth region includes a second object, and the area of the second object in the eroded image is smaller than the area of the second object in the first image; obtaining a first detection region based on the region intersection situation between the fifth region and the third region, where the first detection region is used to indicate the position of the second object in the eroded image; performing a morphological dilation process on the first detection region to obtain a dilated image; performing object screening on the second object in the dilated image to obtain a screened object, where the screened object is used to indicate the second object that occludes the first object; obtaining a first detection box according to the bounding rectangle corresponding to the screened object. Through the above method, by running relevant morphological operations (including erosion operation and dilation operation), the position of the second object in the first image can be accurately located, and the positioning accuracy of the second object can be improved.
[0018] In combination with the first aspect, in some implementations of the first aspect, the above method further includes: obtaining a second detection region based on the region intersection situation between the dilated image and the third region, where the second detection region is used to indicate the position of the screened object in the dilated image; using the second object in the second detection region as the screened object. Through the above method, by selecting the region intersection situation between the dilated image and the second region to determine the screened object, those second objects that do not occlude the region where the first object is located can be filtered out, and the judgment accuracy of the second object can be improved.
[0019] In combination with the first aspect, in some implementations of the first aspect, the above method further includes: performing a first image segmentation process on the first image through a first segmentation model to obtain a first segmented image; performing a second image segmentation process on the first image through a second segmentation model to obtain a second segmented image. Through the above method, by using different segmentation models to perform different image segmentation processes on the first image, two segmented images that divide different regions can be obtained, thereby improving the accuracy of the image segmentation result. Moreover, by using two models to perform different image segmentation operations on the first image respectively, a complex task can be simplified into two lightweight tasks, which are respectively executed by two different image segmentation models, further improving the image processing efficiency and shortening the time delay.
[0020] In combination with the first aspect, in certain implementations of the first aspect, the above method further includes: obtaining a first sample image, where the first sample image includes a first sample object and a second sample object, and the second sample object occludes a first sample area of the first sample object; performing object annotation on the second sample object to obtain an object annotation result of the second sample object in the first sample image, where the object annotation result includes at least one of the object category, position, bounding box, or first mask area of the second sample object in the first sample image; training a second sample model based on the object annotation result and the first sample image to obtain a second segmentation model. Through the above method, training the second sample model based on the method of automatically annotating training data can greatly improve the model training efficiency and accuracy.
[0021] In combination with the first aspect, in certain implementations of the first aspect, the above method further includes: performing category annotation processing on the second sample object to obtain the object category corresponding to the second sample object; performing detection processing on the first sample image based on the object category to obtain a sample detection box of the second sample object in the first sample image, where the sample detection box is used to indicate the position of the second sample object in the first sample image; performing image segmentation processing on the first sample image based on the sample detection box to obtain a first mask area of the second sample object in the first sample image; using the object category, sample detection box, and first mask area as the object annotation result. Through the above method, by using different annotation methods, including category annotation, detection box annotation, and mask annotation, and decomposing the annotation task of training data into multiple different lightweight tasks to be executed separately, the difficulty of the annotation task can be reduced, and the efficiency and accuracy of model training can be further improved.
[0022] In combination with the first aspect, in certain implementations of the first aspect, the above method further includes: obtaining a reference category related to the first sample object; excluding the reference category from the object category to obtain a filtered category; performing detection processing on the first sample image based on the filtered category to obtain a sample detection box of the second sample object in the first sample image. Through the above method, after performing category annotation, excluding the reference category related to the first object from the object category can ensure that the final filtered categories all belong to the categories corresponding to the occluders, improve the accuracy of category localization, and further improve the model training accuracy.
[0023] In combination with the first aspect, in certain implementations of the first aspect, the above method further includes: obtaining a second sample image, where the second sample image includes a third sample object; performing image segmentation processing on the third sample image to obtain a second mask area of the third sample object in the second sample image; training a first sample model based on the second mask area, object annotation result, and second sample image to obtain a first segmentation model.
[0024] In a second aspect, an occlusion determination device is provided. The device includes a unit composed of software and / or hardware for executing any one of the methods in the first aspect.
[0025] In a third aspect, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, any one of the methods in the first aspect can be implemented.
[0026] In a fourth aspect, a chip is provided, including a processor for reading and executing a computer program stored in a memory. When the computer program is executed by the processor, any one of the methods in the first aspect can be implemented.
[0027] Optionally, the chip further includes a memory electrically connected to the processor.
[0028] Optionally, the chip may further include a communication interface.
[0029] In a fifth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, any one of the methods in the first aspect can be implemented.
[0030] In a sixth aspect, a computer program product is provided. The computer program product includes a computer program. When the computer program is executed by a processor, any one of the methods in the first aspect can be implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 is a schematic diagram of occlusion determination in an embodiment of the present application.
[0032] Figure 2 is a schematic diagram of an occlusion determination method in a photo-editing scenario in an embodiment of the present application.
[0033] Figure 3 is a flowchart of an occlusion determination method in an embodiment of the present application.
[0034] Figure 4 is a schematic diagram of a first segmented image in an embodiment of the present application.
[0035] Figure 5 is a schematic diagram of a second segmented image in an embodiment of the present application.
[0036] Figure 6 is a schematic diagram of a process for determining a first detection area in an embodiment of the present application.
[0037] Figure 7 is a schematic diagram of a process for obtaining a second detection area in an embodiment of the present application.
[0038] Figure 8 It is a flowchart of a method for training a segmentation model according to an embodiment of the present application.
[0039] Figure 9 It is a flowchart of a process for generating an object annotation result according to an embodiment of the present application.
[0040] Figure 10 It is a schematic diagram of masking area annotation display according to an embodiment of the present application.
[0041] Figure 11 It is a schematic flowchart of a method for judging occlusion according to an embodiment of the present application.
[0042] Figure 12 It is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application. Detailed implementation manners
[0043] The solutions of the embodiments of the present application will be introduced below with reference to the accompanying drawings.
[0044] In the related solutions, all objects in an image are annotated by extracting various visual features (such as texture, color, shape, etc.) in the image to obtain the annotation results of all objects, and then the occlusion situation of a specified object is judged according to the annotation results. However, since there are too many object categories included in the image, the method of annotating all objects in the image will cause the image processing time to be too long and reduce the image processing efficiency.
[0045] Therefore, how to accurately judge the occlusion situation becomes a key issue.
[0046] In view of this, the present application provides an image processing method. For the sake of easy understanding, a specific example will be used for illustration below. Schematically, please refer to Figure 1 , which shows a schematic diagram of occlusion judgment provided by an exemplary embodiment of the present application. As Figure 1 shown, the first image is currently displayed, where the first image includes a first object and a second object. The first object is in area 110, and the second object is in area 120. As can be seen from Figure 1 , there is an overlapping situation between the area 120 where the second object is located and the actual area where the first object is located. That is to say, area 110 is used to represent the area where the first object is displayed in the first image.
[0047] After performing first image segmentation processing on the first image, a first segmented image 100 is obtained. Among them, in the first segmented image 100, there are area 110 and area 130. Among them, in the first segmented image 100, area 130 is recognized as the area belonging to the first object. Therefore, the actual area corresponding to the first object in the first segmented image 100 is the area sum of area 110 and area 130.
[0048] It should be noted that the above-mentioned area 130 and area 120 are the same area. In some other embodiments, area 130 is smaller than area 120, and the present application does not limit this.
[0049] According to the intersection situation of the area between area 120 and the area, it is determined whether the second object in the first image occludes the first object. As Figure 1 shown, the second object occludes the first object, and the occlusion area is area 140.
[0050] Next, a detailed description will be given for the image retouching scenario. Please refer to Figure 2 , which shows a schematic diagram of the occlusion determination method in the image retouching scenario provided by an exemplary embodiment of the present application. As Figure 2 shown, during the operation of the gallery application on the electronic device, the gallery interface 200 is displayed. Multiple images are displayed in the gallery interface 200, including the first image 201 and the second image 202. Object A is displayed in both the first image 201 and the second image 202. In another case, the first image 201 is displayed in the gallery interface 200, but the second image 202 is stored in the gallery application but not displayed in the gallery interface 200.
[0051] After the first image 201 is selected, the image display interface 210 is displayed. The first image 201 is displayed in the image display interface 210. In addition, an editing control 21 is also displayed in the image display interface 210. When a trigger operation on the editing control 21 is received, the editing interface 230 corresponding to the first image 201 is displayed. At this time, in a realizable case, an occlusion determination is performed on the first image 201 to determine whether object A in the first image 201 is occluded by other objects. Specifically, according to the Figure 1 occlusion determination method described above, if object A is occluded by other objects, the content "the current image cannot be repaired" is displayed in the editing interface 230 ( Figure 2 not displayed in). In addition, in addition to the first image 201, the above-mentioned occlusion determination method is also performed on the second image 202. If the result of the occlusion determination of the second image 202 is that object A is occluded by other objects, the second image 202 is not used as a reference image for the first image 201.
[0052] When the occlusion judgment result corresponding to the first image 201 is that object A is not occluded by other objects, a repair control 22 is displayed in the editing interface 230. When a trigger operation on the repair control 22 is received, a repair editing interface 240 is displayed. In the repair editing interface 240, a closed-eye repair control 23 is displayed, which is used to repair object A in the closed-eye state to the open-eye state. When a trigger operation on the closed-eye repair control 23 is received, in a feasible case, an occlusion judgment is performed on the first image 201 to determine whether object A in the first image 201 is occluded by other objects. Specifically, it can be based on, for example, Figure 1 the occlusion judgment method described above. If object A is occluded by other objects, the content "The current image cannot be repaired" is displayed in the editing interface 230 ( Figure 2 not displayed in Figure 2 ). If the eyes of object A are recognized as not being occluded by other objects, the eye state of object A in the first image 201 is repaired, and finally a third image 203 is obtained. In the third image 203, object A is displayed in the open-eye state.
[0053] That is, according to Figure 2 the content described above, in the image editing scenario, the process of judging occluders for the selected image can be executed at multiple times during the entire image processing process. This application does not specifically limit the execution timing of the occlusion judgment.
[0054] Optionally, the above only describes the state adjustment for one object in the image. In the embodiments of this application, the state of the facial features of a single object in the same image can be adjusted, or the states of the facial features corresponding to multiple objects in the same image can be adjusted. In this case, the same second image can be used (at this time, all objects included in the first image exist in the second image) to adjust the states of the facial features of each object in the first image, or for different objects in the first image, different second images can be used respectively (for example: the first image includes object 1 and object 2, image 1 containing object 1 is used as the second image corresponding to object 1 in the first image, and image 2 containing object 2 is used as the second image corresponding to object 2 in the first image) to adjust the states of the facial features of each object in the first image.
[0055] In addition to the image editing scenario, occlusion judgment can also be applied to the vehicle occlusion judgment scenario. During the driving of the vehicle, when the camera installed on the vehicle captures a road image, the road image is used as the first image. Taking the first image including vehicle a and road b as an example, the recognition task is whether vehicle a occludes the turning of road b. Through the above Figure 1In the recognition method, if it is recognized that vehicle a blocks the turning at the end of road b, the occluded image is further repaired to obtain a final repaired image in which road b has no occlusion, facilitating the user to familiarize with the road conditions.
[0056] Before elaborating on the method provided in the embodiments of the present application, the execution subject provided in the embodiments of the present application will be introduced first. In one case, the method provided in the embodiments of the present application can be executed by an electronic device with an image capture function. The main application scenario can be a scenario of taking pictures through the electronic device. Specifically, it is applied to the image processing scenario of the electronic device. As an example but not limitation, the electronic device can be, but is not limited to, terminals such as mobile phones, smart watches, and tablets. The embodiments of the present application do not limit this. In another case, the method provided in the embodiments of the present application can be executed collaboratively by an electronic device with an image capture function and a server. More specifically, the technical solutions in the embodiments of the present application can be applied to terminal devices and servers.
[0057] The electronic device (i.e., the above-mentioned terminal device) in the embodiments of the present application can be a television, a desktop computer, a laptop computer, or a portable electronic device such as a mobile phone, a tablet computer, a camera, a video camera, a video recorder, or other electronic devices with a photographing function, an electronic device in a 5G network, or an electronic device in a future evolved public land mobile network (PLMN). The present application does not limit this.
[0058] The server in the embodiments of the present application can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Or, the server can also be implemented as a node in a blockchain system. Therefore, the present application does not limit the implementation architecture of the server.
[0059] The occlusion judgment method provided in this embodiment can be executed independently by the electronic device, independently by the server, or collaboratively by the electronic device and the service. The present embodiment does not limit this.
[0060] It should be noted that the information involved in this application (including but not limited to object device information, object personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals are all authorized by the object or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards in the relevant regions.
[0061] Next, a detailed description will be given of the reasoning process of the occlusion judgment scheme.
[0062] Schematically, please refer to Figure 3 , which shows a flowchart of an occlusion judgment method provided by an exemplary embodiment of the present application. As Figure 3 shown, the method includes the following steps.
[0063] Step 310, obtain a first image.
[0064] Obtain a first image, where the first image includes a first object and a second object. Based on the first object, in this embodiment, it is determined whether the second object occludes the first object.
[0065] Step 320, input the first segmentation model.
[0066] During the process of inputting the first image into the first segmentation model, the first segmentation model performs pixel point prediction on the second region where the second object is located. If the prediction result is that the second object does not occlude the first object, the region where the second object is located is identified as the background region. If the prediction result is that the second object occludes the first object, the region corresponding to the second object is identified as the foreground region. At this time, the region where the first object is located is also identified as the foreground region. If the region where the second object is located includes region a and region b, where region a occludes the first object and region b does not occlude, then region a is identified as the foreground region and region b is identified as the background region. Through this first segmentation model, the region where the first object is occluded by the second object can also be identified as the foreground region of the same type as the region where the first object is located, so that it ignores the regional impact of the occluder on the first object.
[0067] It should be noted that for the first segmentation model, the object of concern is the foreground region.
[0068] Step 330, input the second segmentation model.
[0069] Perform image segmentation on the first image through the second segmentation model, take the area where the first object is shown in the first image as the foreground area, and identify all other areas except the first object as the background area. That is to say, if the second object occludes the first object and the occluded area is area c, then identify area c as the background area. If the second object does not occlude the first object and the area where the second object is located in the first image is area d, then also identify area d as the background area.
[0070] That is, the second segmentation model is used to segment the foreground area and the background area of the first image.
[0071] That is to say, the difference between the first segmentation model and the second segmentation model is that the first segmentation model can specifically identify the occluding situation of the second object to the first object as the foreground area or the background area. In the second segmentation model, regardless of whether the second object occludes the first object, the area where the second object is shown in the first image is the background area.
[0072] In this embodiment, for the second segmentation model, the concerned object is the background area.
[0073] Step 340, obtain the first segmentation image.
[0074] Schematically, when performing image segmentation on the first image through the first segmentation model, in the case where the second object occludes the first object, the occluded area is the fourth area, and the second object is in the second area in the first image, then the first segmentation image includes the first area where the first object is shown in the first image, and the fourth area where the first object is occluded by the second object. Take the first area and the fourth area as the third area, that is, the foreground area corresponding to the first image, and take the area other than the foreground area in the first image as the background area.
[0075] Schematically, please refer to Figure 4 , which shows a schematic diagram of the first segmentation image provided by an exemplary embodiment of the present application. As Figure 4 shown, currently the first segmentation image 400 is displayed, which includes a white area 420 and a black area 410. Among them, the white area 410 refers to the foreground area. If there is a second object occluding the first object in the foreground area, also identify the occluded area as the foreground area. The black area 420 refers to the background area.
[0076] Step 350, obtain the second segmentation image.
[0077] Schematically, when performing segmentation on the first image through the second segmentation model, take the area where the first object is shown in the first image as the foreground area, and identify the other areas outside the foreground area as the background area.
[0078] Schematic. Please refer to Figure 5 , which shows a schematic diagram of a second segmented image provided by an exemplary embodiment of the present application. As Figure 5 shown, the second segmented image 500 is currently displayed, which includes a white area 510 and a black area 520. Among them, the white area 510 refers to the foreground area, and the black area 520 refers to the background area.
[0079] Step 360, obtain the erosion result.
[0080] Perform morphological erosion processing on the second segmented image to obtain the erosion result corresponding to the first segmented image.
[0081] Among them, morphological erosion processing is an operation based on a structuring element (also called an erosion kernel). The structuring element is a small-sized shape template used to compare pixel by pixel with the second segmented image. During the erosion process, the structuring element is placed at a certain pixel position in the first segmented image. If all pixels within the structuring element match the corresponding pixels in the image (usually in a binary image, matching means both are 1 or both are 0, depending on the definition of erosion and the background value), then the pixel remains unchanged; otherwise, the pixel is set to the background value (usually 0). This process traverses every pixel in the image until the entire second segmented image is processed.
[0082] In this embodiment, the structuring element is a preset value. For example, the kernel size is an elliptical structure of 5*5.
[0083] When the erosion result is obtained, the area where the second object is displayed in the erosion result (i.e., the fifth area) will be slightly smaller than the area corresponding to the second object in the second segmented image. In this way, during the execution of step 370, the background area in the erosion result will not intersect with the foreground area in the second segmented image, and the actual position of the second object in the first image can be more accurately located.
[0084] Step 370, determine the first detection area.
[0085] Schematic. Calculate the area intersection situation between the foreground area in the first segmented image and the background area in the erosion result, and determine the specified area that belongs to both the foreground area in the first segmented image and the background area in the dilation result as the first detection area.
[0086] Schematic. Please refer to Figure 6, which shows a schematic diagram of the first detection area determination process provided by an exemplary embodiment of the present application. Among them, the first image is subjected to image segmentation processing through the above-mentioned first segmentation model and the second segmentation model to obtain a first segmented image and a second segmented image. Among them, after the second segmented image is subjected to morphological erosion processing, an eroded image is obtained. Therefore, in the eroded image, a foreground area 610 and a background area 620 (the slanted area indicates the background area) are included. In the first segmented image, a foreground area 630 (the slanted line indicates the foreground area) is included. According to the intersection situation of the area between the background area 620 and the foreground area 630, the area 640 is determined as the first detection area, where the first detection area is the area corresponding to the second object in the first image after morphological erosion.
[0087] Step 380, obtain the dilated image.
[0088] Since the first detection area is the area corresponding to the second object in the first image after morphological erosion, it is also necessary to perform a morphological dilation operation on the image corresponding to the first detection area to restore the area corresponding to the second object in the first detection area to the actual area in the first image.
[0089] Step 390, determine the second detection area.
[0090] After obtaining the dilated image, according to the intersection situation of the area between the dilated image and the first segmented image, the areas that do not belong to the area where the first object is located but are also recognized as the first detection area can be excluded, so that all the screened objects in the final second detection area have an intersection with the area where the first object is located, which helps to exclude other objects that do not block the first object.
[0091] Schematically, please refer to Figure 7 , which shows a schematic diagram of the second detection area acquisition process provided by an exemplary embodiment of the present application. After obtaining the dilated image, the dilated image includes multiple first detection areas, including area 701, area 702, and area 703. According to the intersection situation of the area between the dilated image and the first segmented image, it is finally determined that area 701 is an area that belongs to both the first detection area and has an intersection with the area where the first object is located. Therefore, area 701 is determined as the second detection area.
[0092] Step 3100, obtain the first detection box.
[0093] After determining the second detection area, according to the area contour corresponding to the second detection area, the minimum bounding rectangle corresponding to the second detection area is determined and used as the first detection box.
[0094] Step 3110, obtain the eye rectangle box.
[0095] In this embodiment, when the first image is obtained, for the first object in the first image, based on the eye region of the first object, it is determined whether the second object occludes the eye region of the first object. Therefore, after the first image is obtained, the first image needs to be cropped to obtain the eye rectangular frame corresponding to the first image.
[0096] Step 3120, obtain the face rectangular frame.
[0097] In this embodiment, when the first image is obtained, for the first object in the first image, based on the face region of the first object, it is determined whether the second object occludes the face region of the first object. Therefore, after the first image is obtained, the first image needs to be cropped to obtain the face rectangular frame corresponding to the first image.
[0098] Obtain the intersection relationship between the face rectangular frame / eye rectangular frame and the first detection frame. If there is an intersection between the face rectangular frame / eye rectangular frame and the first detection frame, then execute step 3130; otherwise, execute step 3140.
[0099] Step 3130, occlusion result.
[0100] If there is an intersection between the face rectangular frame / eye rectangular frame and the first detection frame, it indicates that the second object occludes the face region / eye region of the first object.
[0101] Step 3140, non-occlusion result.
[0102] If there is no intersection between the face rectangular frame / eye rectangular frame and the first detection frame, it indicates that the second object does not occlude the face region / eye region of the first object.
[0103] In this embodiment, the intersection over union (IoU) with the face rectangular frame / eye rectangular frame is calculated. If the IoU is greater than 0, it is considered that the face region / eye region is occluded; if it is equal to 0, it is considered that the face region / eye region is not occluded.
[0104] Next, the training process of the occlusion solution will be described in detail.
[0105] Schematically, please refer to Figure 8 , which shows a flowchart of a segmentation model training method provided by an exemplary embodiment of the present application. The method includes the following steps.
[0106] Step 810, obtain the first sample image.
[0107] Schematically, the first sample image is obtained from the training dataset. The first sample image includes the first sample image and the second sample image, where the second sample image occludes the first sample region corresponding to the first sample image.
[0108] Step 820, object annotation.
[0109] Perform object annotation on the second sample object in the first sample image to obtain the object annotation result of the second sample object in the first sample image.
[0110] Step 830, obtain the object annotation result.
[0111] Optionally, the object annotation result includes at least one of the object category, position, bounding box, or first mask region of the second sample object in the first sample image.
[0112] Next, the process of obtaining the object annotation result will be described in detail.
[0113] Schematically, please refer to Figure 9 , which shows a flowchart of the object annotation result generation process provided by an exemplary embodiment of the present application. As Figure 9 shown, the method includes the following steps.
[0114] Step 910, obtain the first sample image.
[0115] Schematically, the first sample image is obtained from the training dataset. The first sample image includes the first sample image and the second sample image, where the second sample image occludes the first sample region corresponding to the first sample image.
[0116] Step 920, category tagging.
[0117] Perform category annotation processing on the second sample object to obtain the object category corresponding to the second sample object.
[0118] In this embodiment, the recognize anything model (RAM model) is called to perform object tagging on all the second sample objects that occlude the first sample object in the first sample image, and the object categories corresponding to each of the second sample objects are obtained. Among them, if there are two identical second sample objects, a set of second sample objects is generated, and the corresponding object category is annotated for this set of sample objects.
[0119] Step 930, category elimination.
[0120] After obtaining the object categories corresponding to each of the second sample objects, first obtain the reference categories related to the first object: for example, eyes, nose, mouth, etc. If the above reference categories exist in the multiple object categories, these reference categories are eliminated from the object categories.
[0121] Step 940, filter categories.
[0122] If there are the above reference categories among multiple object categories, after removing these reference categories from the object categories, the remaining object categories are used as the screening categories.
[0123] Step 950: Obtain the sample detection box.
[0124] Perform bounding box detection on the screening categories through the detection model (Grounding DINO) to determine the sample detection box (bounding box) of the second sample object corresponding to the screening categories in the first sample image. The sample detection box is used to determine the position of the second sample object in the first sample image.
[0125] Step 960: Obtain the first mask region.
[0126] Input the sample detection box into the trained segmentation model (segment anything model, SAM model) to generate the first mask region corresponding to the sample detection box.
[0127] Step 970: Save.
[0128] Store the first mask region of the second sample object corresponding to the screening categories as annotation data.
[0129] Schematically, please refer to Figure 10 , which shows a schematic diagram of mask region annotation display according to an exemplary embodiment of the present application. As Figure 10 shown, the first sample image is currently displayed. The first sample image includes the first sample object and the second sample object that occludes the first sample object. Through the above object annotation method, the first mask region 1010 corresponding to the second sample object is finally obtained, and also, the object category corresponding to the second sample object: pen.
[0130] Step 840: Obtain the second sample image.
[0131] Schematically, the second sample image is obtained from the training dataset. Among them, the second sample image includes the third sample object, but the third sample object is not occluded by other objects.
[0132] Step 850: Obtain the second mask region.
[0133] Schematically, the second mask region is used to indicate the region corresponding to the specified part features (e.g., facial features) of the third sample object in the second sample image.
[0134] Optionally, the second mask region is the data that has been labeled in the training dataset; or, a pre-trained mask recognition model is used to perform mask recognition on the second sample image to determine the second mask region corresponding to the third sample object.
[0135] Step 860: Train to obtain the second segmentation model.
[0136] Train the second sample model according to the object annotation result and the first sample image to obtain the second segmentation model.
[0137] Step 870: Train to obtain the first segmentation model.
[0138] Train the first sample model according to the second mask region, the object annotation result corresponding to the first sample image, and the second sample image, and finally obtain the first segmentation model.
[0139] It should be noted that regarding the training order of the first segmentation model and the second segmentation model, the first segmentation model can be trained first, and then the second segmentation model can be trained; or the second segmentation model can be trained first, and then the first segmentation model can be trained; or the first segmentation model and the second segmentation model can be trained synchronously. The embodiments of the present application do not limit this.
[0140] Figure 11 It is a schematic flowchart of a method for judging occlusion in an embodiment of the present application. The following Figure 11 introduces each step shown. Figure 11 It can be executed by Figures 1 - 10 any one of the electronic devices in.
[0141] S1001: Obtain the first image.
[0142] Among them, the first image includes a first object and a second object. The first object is in a first region, and the second object is in a second region.
[0143] It should be understood that the first object and the second object in the first image are only used to distinguish different objects. According to the selected objects, the corresponding occlusion situations are also different. For example: if the first object is object a and the second object is a fan, if object a is used as a reference, then when the fan blocks the face of object a, it can be considered that the second object occludes the first object. If the fan is used as a reference, then the fan blocks the face of object a, but the face of object a does not occlude the fan. However, since the hand of object a holds the handle of the fan, therefore, the hand of object a occludes the fan.
[0144] In this embodiment, object a is used as a reference to determine whether the second object occludes the first object.
[0145] S1002. Perform a first image segmentation process on the first image to obtain a first segmented image, where the first segmented image includes a third region corresponding to the first image.
[0146] Among them, the third region includes the first region, and in the case where the first object is blocked by the second object, a fourth region obtained after performing an object repair process on the first object.
[0147] Schematically, in the first image, the first object is in the first region, and the second object is in the second region. Among them, the first region and the second region are regions that do not intersect in the first image.
[0148] Optionally, the first object is completely in the second region, or part of the first object is in the first region and part is in the second region (in this case, part of the region is blocked by the second object). In this embodiment, the first region and the second region are determined only based on the display situation of the first object and the second object in the first image. That is to say, if there is a region a in the first object that is blocked by the second object, then region a will belong to the second region.
[0149] Schematically, the third region includes the first region and the fourth region. As mentioned above, if there is a region a in the first object that is blocked by the second object, then during the first image segmentation process of the first image, it is recognized that region a is blocked by the second object. Ignoring the blocking situation of the second object on region a, the other regions of the first object except region a are used as the first region, region a is used as the fourth region, and the first region and the fourth region are used as the third region.
[0150] S1003. Determine whether the second object blocks the first object according to the region intersection situation between the second region and the third region.
[0151] Among them, the region intersection situation is used to indicate the overlapping region formed after the bounding box of the second region overlaps with the bounding box of the third region.
[0152] It should be understood that the third region is used to indicate the complete region where the first object is located in the first image when there is no occlusion of the first object. The first region may be the same as the third region, or the first region may be smaller than the third region (that is, at this time, part of the first object is blocked by the second object).
[0153] Schematically, the region intersection situation refers to the overlapping situation between the bounding box corresponding to the second region and the bounding box corresponding to the third region. If there is an overlapping region between the bounding box of the second region and the bounding box of the third region, it means that the second object blocks the first object. If there is no overlapping region between the bounding box of the second region and the bounding box of the third region, it means that the second object does not block the first object.
[0154] In some implementations, the above method further includes: obtaining a reference detection box of a first object in a first image, where the reference detection box is used to indicate the position of the first object in the first image; obtaining a first detection box of a second object in the first image according to the area intersection between a second region and a third region, where the first detection box is used to indicate the position of the second object in the first image; determining whether the second object occludes the first object according to the overlap between the first detection box and the reference detection box. By the above method, a reference detection box corresponding to the first object in the first image is set, and after obtaining the first detection box of the second object in the first image according to the area overlap between the second region and the third region, the edge overlap degree between the reference detection box and the first detection box is calculated, thereby improving the accuracy of occlusion judgment.
[0155] In some implementations, the above method further includes: when there is an overlapping area between the first detection box and the reference detection box, determining the occlusion area of the second object on the first object according to the area of the overlapping area. By the above method, when it is determined that the second object occludes the first object, the occlusion area of the second object on the first object can also be determined according to the area of the overlapping area between the first detection box and the reference detection box. In addition to judging whether there is occlusion, the occlusion area can also be judged, further improving the accuracy of the occlusion judgment result.
[0156] In some implementations, the above method further includes: performing a second image segmentation process on the first image to obtain a second segmented image, where the second segmented image includes a first region and a second region; performing a morphological erosion process on the second segmented image to obtain an erosion image, where the erosion image includes a fifth region, and the fifth region includes the second object, and the area of the second object in the erosion image is smaller than the area of the second object in the first image; obtaining a first detection region according to the area intersection between the fifth region and the third region, where the first detection region is used to indicate the position of the second object in the erosion image; performing a morphological dilation process on the first detection region to obtain a dilation image; performing object screening on the second object in the dilation image to obtain a screened object, where the screened object is used to indicate the second object that occludes the first object; obtaining the first detection box according to the circumscribed rectangle corresponding to the screened object. By the above method, related operations of morphology (including erosion operation and dilation operation) are performed to accurately locate the position of the second object in the first image, improving the positioning accuracy of the second object.
[0157] In some implementations, the above method further includes: obtaining a second detection region based on the region intersection between the dilated image and the third region, where the second detection region is used to indicate the position of the screening object in the dilated image; and taking the second object within the second detection region as the screening object. By determining the screening object based on the region intersection between the dilated image and the second region in the above manner, those second objects that do not occlude the region where the first object is located can be screened out, improving the judgment accuracy of the second object.
[0158] In some implementations, the above method further includes: performing a first image segmentation process on the first image through a first segmentation model to obtain a first segmented image; and performing a second image segmentation process on the first image through a second segmentation model to obtain a second segmented image. By the above method, different image segmentation processes are performed on the first image using different segmentation models, so as to obtain two segmented images that divide different regions respectively, improving the accuracy of the image segmentation result. Moreover, by using two models to perform different image segmentation operations on the first image respectively, a complex task can be simplified into two lightweight tasks, which are respectively executed by two different image segmentation models, further improving the image processing efficiency and shortening the time delay.
[0159] In some implementations, the above method further includes: obtaining a first sample image, which includes a first sample object and a second sample object, where the second sample object occludes the first sample region of the first sample object; performing object annotation on the second sample object to obtain an object annotation result of the second sample object in the first sample image, where the object annotation result includes at least one of the object category, position, bounding box, or first mask region of the second sample object in the first sample image; and training a second sample model based on the object annotation result and the first sample image to obtain a second segmentation model. By the above method, training the second sample model based on the method of automatically annotating training data can greatly improve the model training efficiency and accuracy.
[0160] In some implementations, the above method further includes: performing a category annotation process on the second sample object to obtain the object category corresponding to the second sample object; performing a detection process on the first sample image based on the object category to obtain a sample detection box of the second sample object in the first sample image, where the sample detection box is used to indicate the position of the second sample object in the first sample image; performing an image segmentation process on the first sample image based on the sample detection box to obtain a first mask region of the second sample object in the first sample image; using the object category, the sample detection box, and the first mask region as the object annotation result. Through the above method, by using different annotation methods, including category annotation, detection box annotation, and mask annotation, the annotation task of the training data is decomposed into multiple different lightweight tasks to be executed separately, which can reduce the difficulty of the annotation task and further improve the efficiency and accuracy of model training.
[0161] In some implementations, the above method further includes: obtaining a reference category related to the first sample object; excluding the reference category from the object category to obtain a filtered category; performing a detection process on the first sample image based on the filtered category to obtain a sample detection box of the second sample object in the first sample image. In the above manner, after performing the category annotation, excluding the reference category related to the first object from the object category can ensure that the final filtered categories all belong to the categories corresponding to the occluders, improve the accuracy of category localization, and further improve the model training accuracy.
[0162] In some implementations, the above method further includes: obtaining a second sample image, where the second sample image includes a third sample object; performing an image segmentation process on the third sample image to obtain a second mask region of the third sample object in the second sample image; training the first sample model based on the second mask region, the object annotation result, and the second sample image to obtain a first segmentation model. In the technical solution of the present application, after obtaining the first image including the first object and the second object, when the first object and the second object are respectively in the first region and the second region, performing a first segmentation process on the first image to obtain a first segmentation image, where the first segmentation image shows a third region corresponding to the first image, and the third region not only includes the first region but also includes a fourth region after performing an object repair process when the first object is occluded. Determining whether the second object occludes the first object according to the region intersection between the second region and the third region. That is, on the basis of segmenting the first image into foreground and background images, a first segmentation process is also performed on the first image, and an object repair function is automatically executed on the occluded first object in the first segmentation image, so that the third region corresponding to the first object can be completely reflected in the first segmentation image. Thus, according to the region intersection between the second region including only the second object and the first region including only the first object, finally determining whether the second object causes occlusion to the first object can improve the accuracy of the occlusion judgment result.
[0163] The above mainly introduced the method of the embodiments of the present application in combination with the accompanying drawings. It should be understood that although the steps in the flowcharts involved in the above-mentioned embodiments are shown in sequence, these steps are not necessarily executed in the order shown in the figures. Unless there is a clear description in this article, there is no strict order limit for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps. The device of the embodiments of the present application will be introduced below in combination with the accompanying drawings.
[0164] Figure 12 is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application. As Figure 12 shown, the electronic device 1200 may include a processor 1210, an external memory interface 1220, an internal memory 1221, a universal serial bus (USB) interface 1230, a charging management module 1240, a power management module 1241, a battery 1242, an antenna 1, an antenna 2, a mobile communication module 1250, a wireless communication module 1260, an audio module 1270, a speaker 1270A, a receiver 1270B, a microphone 1270C, a headset interface 1270D, a sensor module 1280, a button 1290, a motor 991, an indicator 992, a camera 993, a display screen 994, and a subscriber identification module (SIM) card interface 995, etc.
[0165] Among them, the sensor module 1280 may include a pressure sensor 1280A, a gyroscope sensor 1280B, a barometric pressure sensor 1280C, a magnetic sensor 1280D, an acceleration sensor 1280E, a distance sensor 1280F, a proximity light sensor 1280G, a fingerprint sensor 1280H, a temperature sensor 1280J, a touch sensor 1280K, an ambient light sensor 1280L, a bone conduction sensor 1280M, etc.
[0166] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device 1200. In other embodiments of the present application, the electronic device 1200 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0167] Exemplarily, Figure 12 The illustrated processor 1210 may include one or more processing units. For example, the processor 910 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0168] Among them, the controller may be the nerve center and command center of the electronic device 900. The controller can generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching and executing instructions.
[0169] In the embodiments of the present application, the processor 210 is mainly used to start the application loading page and respond to the interactive operations input by the user.
[0170] A memory may also be provided in the processor 1210 for storing instructions and data. In some embodiments, the memory in the processor 1210 is a cache memory. This memory can save the instructions or data that the processor 1210 has just used or recycled. If the processor 1210 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 1210, and thus improves the efficiency of the system.
[0171] The electronic device 1200 realizes the display function through the GPU, the display screen 994, and the application processor, etc. The GPU is a microprocessor for image processing, connecting the display screen 994 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 1210 may include one or more GPUs, which execute program instructions to generate or change display information.
[0172] The display screen 994 is used to display images, videos, etc. The display screen 994 includes a display panel. In some embodiments, the electronic device 1200 may include one or N display screens 994, where N is a positive integer greater than 1.
[0173] In the embodiments of the present application, various interfaces are mainly presented to the user through the display screen 994. For example Figure 1 each interface in
[0174] The NPU is a neural-network (NN) computing processor. By borrowing the structure of a biological neural network, for example, borrowing the transmission mode between human brain neurons, it can quickly process input information and can also continuously self-learn. Through the NPU, applications such as intelligent cognition of the electronic device 1200 can be realized, such as: image recognition, face recognition, speech recognition, text understanding, etc.
[0175] In the embodiments of the present application, the NPU can be used to recognize the displayed content and instruct the processor 1210 to generate the displayed content. For example, it can instruct the processor 1210 to generate the displayed content in the Figure 1 form shown.
[0176] The pressure sensor 1280A is used to sense pressure signals and can convert the pressure signals into electrical signals. In some embodiments, the pressure sensor 1280A can be disposed on the display screen 994. There are many types of pressure sensors 1280A, such as resistive pressure sensors, inductive pressure sensors, capacitive pressure sensors, etc. The capacitive pressure sensor can include at least two parallel plates with conductive materials. When a force acts on the pressure sensor 1280A, the capacitance between the electrodes changes. The electronic device 1200 determines the intensity of the pressure according to the change in capacitance. When a touch operation acts on the display screen 994, the electronic device 1200 detects the intensity of the touch operation according to the pressure sensor 1280A. The electronic device 1200 can also calculate the position of the touch according to the detection signal of the pressure sensor 1280A. In some embodiments, touch operations with the same touch position but different touch operation intensities can correspond to different operation instructions. For example, when a touch operation with a touch operation intensity less than the first pressure threshold acts on the short message application icon, the instruction to view the short message is executed. When a touch operation with a touch operation intensity greater than or equal to the first pressure threshold acts on the short message application icon, the instruction to create a new short message is executed.
[0177] In the embodiments of the present application, the pressure sensor 1280A is mainly used to collect the user's interaction operations, such as sliding operations, click operations, etc.
[0178] It should be noted that for the information interaction, execution process, etc. between the above-mentioned devices / units, since they are based on the same concept as the method embodiments of this application, for their specific functions and the technical effects brought, reference can be made to the method embodiment part, and details will not be repeated here.
[0179] Those skilled in the art can clearly understand that for the convenience and conciseness of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiments, and details will not be repeated here.
[0180] The embodiment of this application also provides an electronic device, which includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor. When the processor executes the computer program, the steps in any of the above methods can be implemented.
[0181] The embodiment of this application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps in each of the above method embodiments can be implemented.
[0182] The embodiment of this application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in each of the above method embodiments can be implemented.
[0183] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of this application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the photographing device / electronic device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.
[0184] In the above embodiments, the descriptions of each embodiment have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0185] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0186] In the embodiments provided in this application, it should be understood that the disclosed device / equipment and method can be implemented in other ways. For example, the device / equipment embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical, or other form.
[0187] The unit described as a separation component may or may not be physically separated, and the component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0188] It should be understood that when used in the specification and appended claims of this application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.
[0189] It should also be understood that the term "and / or" used in the specification and appended claims of this application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0190] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" according to the context.
[0191] In addition, in the description of the specification and appended claims of this application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0192] The reference to "one embodiment" or "some embodiments" etc. described in the specification of this application means that a specific feature, structure or characteristic described in combination with the embodiment is included in one or more embodiments of this application. Thus, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized.
[0193] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included within the protection scope of the present application.
Claims
1. A method for determining occlusion, characterized in that: The method comprises: Acquire a first image, where the first image includes a first object and a second object, the first object is in a first area, and the second object is in a second area; performing first image segmentation processing on the first image to obtain a first segmented image, wherein the first segmented image includes a third region corresponding to the first image, the third region including the first region, and a fourth region obtained by performing object restoration processing on the first object when the first object is occluded by the second object; Whether the second object obscures the first object is determined based on an area intersection between the second area and the third area, where the area intersection indicates an overlapping area formed by overlapping bounding boxes of the second area and the third area.
2. The method according to claim 1, characterized in that The determining, based on an area intersection between the second area and the third area, whether the second object obscures the first object includes: Acquire a reference detection frame of the first object in the first image, where the reference detection frame is used to indicate a position of the first object in the first image; Obtaining a first detection frame of the second object in the first image according to an area intersection between the second area and the third area, where the first detection frame is used to indicate a position of the second object in the first image; Determine whether the second object occludes the first object based on an overlap between the first detection frame and the reference detection frame.
3. The method according to claim 2, characterized in that The method further comprises: In a case where there is an overlapping area between the first detection frame and the reference detection frame, the occlusion area of the first object caused by the second object is determined according to the area of the overlapping area.
4. The method according to claim 2, characterized in that The obtaining, based on the intersection of the second area and the third area, that the second object is located before the first detection frame in the first image further includes: performing second image segmentation processing on the first image to obtain a second segmented image, where the second segmented image includes the first region and the second region; The obtaining, based on an area intersection between the first area and the third area, a first detection frame of the second object in the first image includes: performing morphological erosion processing on the second segmented image to obtain an eroded image, wherein the eroded image includes a fifth region, the fifth region includes the second object, and an area of the second object in the eroded image is smaller than an area of the second object in the first image; Obtaining a first detection area based on an intersection between the fifth area and the third area, where the first detection area is used to indicate a position of the second object in the eroded image; performing morphological dilation processing on the first detection area to obtain a dilated image; performing object screening on the second object in the dilated image to obtain a screened object, where the screened object is used to indicate the second object that blocks the first object; The first detection frame is obtained according to the circumscribed rectangular frame corresponding to the screening object.
5. The method according to claim 4, characterized in that The performing object screening on the second object in the dilated image to obtain the screened object includes: obtaining a second detection area based on an intersection between the dilated image and the third area, wherein the second detection area is used to indicate a position of the screening object in the dilated image; A second object in the second detection area is used as the screening object.
6. The method according to claim 4, characterized in that The performing first image segmentation processing on the first image to obtain a first segmented image includes: Performing first image segmentation processing on the first image using a first segmentation model to obtain the first segmented image; The performing second image segmentation processing on the first image to obtain a second segmented image includes: Performing second image segmentation processing on the first image using a second segmentation model to obtain the second segmented image.
7. The method according to claim 6, characterized in that Before performing second image segmentation processing on the first image using the second segmentation model to obtain the second segmented image, the method further includes: Acquire a first sample image, where the first sample image includes a first sample object and a second sample object, and the second sample object covers a first sample area of the first sample object; performing object labeling on the second sample object to obtain an object labeling result of the second sample object in the first sample image, wherein the object labeling result includes at least one of an object category, a position, a bounding box, or a first mask area of the second sample object in the first sample image; A second sample model is trained based on the object labeling result and the first sample image to obtain the second segmentation model.
8. The method according to claim 7, characterized in that The performing object labeling on the second sample object to obtain an object labeling result of the second sample object in the first sample image includes: performing category labeling processing on the second sample object to obtain an object category corresponding to the second sample object; performing detection processing on the first sample image based on the object category to obtain a sample detection frame of the second sample object in the first sample image, where the sample detection frame is used to indicate a position of the second sample object in the first sample image; performing image segmentation processing on the first sample image based on the sample detection frame to obtain a first mask area of the second sample object in the first sample image; The object category, the sample detection frame, and the first mask area are used as the object labeling result.
9. The method according to claim 8, characterized in that The performing detection processing on the first sample image based on the object category to obtain a sample detection frame of the second sample object in the first sample image includes: obtaining a reference category related to the first sample object; Eliminating the reference category from the object category to obtain a screening category; Detection processing is performed on the first sample image based on the screening category to obtain a sample detection frame of the second sample object in the first sample image.
10. The method according to claim 7, characterized in that Before performing first image segmentation processing on the first image using the first segmentation model to obtain the first segmented image, the method further includes: Acquire a second sample image, where the second sample image includes a third sample object; performing image segmentation processing on the second sample image to obtain a second mask region of the third sample object in the second sample image; The first sample model is trained based on the second mask area, the object annotation result and the second sample image to obtain the first segmentation model.
11. An electronic device, characterized in that: The electronic device includes a memory, one or more processors, and a computer program stored in the memory and executable on the processors, wherein when the one or more processors execute the computer program, the electronic device implements the method according to any one of claims 1 to 10.
12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by an electronic device, the method according to any one of claims 1 to 10 is implemented.
13. A computer program product, characterized in that A computer program is included which, when executed, causes the method according to any one of claims 1 to 10 to be performed.
Citation Information
Patent Citations
Image processing method and device, electronic equipment and computer readable storage medium
CN112102340A
Object posture recognition method and device, electronic equipment and storage medium
CN115083021A
Image processing method and electronic equipment
CN115908120A
Vehicle shielding identification method and related device
CN117197796A
Image processing method and electronic equipment
CN118096593A