Image processing method, apparatus, device, and computer-readable storage medium

CN116452651BActive Publication Date: 2026-09-29TCL TECHNOLOGY GROUP CORPORATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210009406.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-06
Publication Date
2026-09-29
Estimated Expiration
2042-01-06

AI Technical Summary

Benefits of technology

[0016]本申请实施例提供的技术方案,获取待处理图像中第一目标物所在的区域和第二目标物所在的区域,计算两个区域之间的交并比,当交并比大于第一预设阈值时,可以判断出第一目标物与第二目标物在同一视线上,在这样的情况下,再计算两个区域在待处理图像中所占比例的差值,根据差值生成第一目标物与第二目标物的位置关系信息。采用本申请实施例的方案,基于两个区域的交并比和两个区域在图像中占比的差值两个维度进行检测,实现了物体之间的位置关系的准确检测。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116452651B_ABST
    Figure CN116452651B_ABST
Patent Text Reader

Abstract

The application discloses an image processing method, device and equipment and a computer readable storage medium. The method comprises the following steps: obtaining a first region where a first target object is located and a second region where a second target object is located in a to-be-processed image; calculating an intersection and union ratio of the first region and the second region; when the intersection and union ratio is greater than a first preset threshold, calculating a difference value of a proportion of the first region and the second region in the to-be-processed image; and generating position relationship information of the first target object and the second target object according to the difference value. By using the application, the position relationship between objects in an image can be accurately detected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, specifically to an image processing method, apparatus, device, and computer-readable storage medium. Background Technology

[0002] In some image-based applications, it is necessary to detect the positional relationships between objects in an image, but current detection methods suffer from low accuracy. Summary of the Invention

[0003] This application provides an image processing method, apparatus, device, and computer-readable storage medium that can improve the accuracy of positional relationship detection between objects.

[0004] In a first aspect, embodiments of this application provide an image processing method, including:

[0005] Obtain the first region where the first target object is located and the second region where the second target object is located in the image to be processed;

[0006] Calculate the intersection-union ratio of the first region and the second region;

[0007] When the crossover ratio is greater than the first preset threshold, the difference between the proportions of the first region and the second region in the image to be processed is calculated.

[0008] The positional relationship information between the first target and the second target is generated based on the difference.

[0009] Secondly, embodiments of this application also provide an image processing apparatus, comprising:

[0010] The target detection module is used to obtain the first region where the first target object is located and the second region where the second target object is located in the image to be processed;

[0011] The first calculation module is used to calculate the intersection-union ratio of the first region and the second region;

[0012] The second calculation module is used to calculate the difference in the proportion of the first region and the second region in the image to be processed when the cross-union ratio is greater than the first preset threshold.

[0013] The image processing module is used to generate positional relationship information between the first target object and the second target object based on the difference.

[0014] Thirdly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in the image processing method provided in any embodiment of this application.

[0015] Fourthly, embodiments of this application also provide an electronic device, which includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps in the image processing method provided in any embodiment of this application.

[0016] The technical solution provided in this application embodiment obtains the regions containing a first target object and a second target object in the image to be processed, calculates the intersection-union ratio (IUGR) between the two regions, and determines that the first and second target objects are on the same line of sight when the IUGR is greater than a first preset threshold. In this case, the difference in the proportions of the two regions in the image to be processed is calculated, and positional relationship information between the first and second target objects is generated based on the difference. By adopting the solution of this application embodiment, detection is performed based on two dimensions: the IUGR of the two regions and the difference in the proportions of the two regions in the image, thus achieving accurate detection of the positional relationship between objects. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram of a first flowchart of an image processing method provided in an embodiment of this application.

[0019] Figure 2 This is a schematic diagram of a target detection scene in the image processing method provided in the embodiments of this application.

[0020] Figure 3 This is a schematic diagram of a scenario where a mobile phone is held in hand.

[0021] Figure 4 This is a schematic diagram of the structure of the image processing apparatus provided in the embodiments of this application.

[0022] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0023] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the protection scope of this application.

[0024] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0025] This application provides an image processing method, which can be executed by an electronic device. The electronic device can be a smartphone, tablet computer, PDA, laptop computer, or desktop computer, etc.

[0026] Please see Figure 1 , Figure 1 This is a schematic diagram of a first flowchart of the image processing method provided in this application embodiment. The specific flow of the image processing method provided in this application embodiment can be as follows:

[0027] 101. Obtain the first region where the target hand is located and the second region where the target object is located in the image to be processed.

[0028] The image to be processed in this embodiment can be a video frame captured from a continuous video frame, or it can be a single image. For example, the electronic device receives an image sent by an external device as the image to be processed. Alternatively, the electronic device captures a picture of a target object in a target space using an image acquisition device, and uses the captured image as the image to be processed.

[0029] For example, in some application scenarios that require gesture recognition, in order to avoid misjudging gestures, the image of the object to be detected can be used as the image to be processed. The image to be processed is first processed using the solution of this application, and then the subsequent gesture recognition is performed based on the image processing results.

[0030] After acquiring the image to be processed, detection can be performed on the image to determine the first region where the first target object is located and the second region where the second target object is located. For example, in one embodiment, target recognition processing is performed on the first and second target objects in the image to be processed according to a target detection model to determine their respective regions.

[0031] Next, taking the target hand as an example, the solution of this application embodiment will be described in detail. For example, this application embodiment detects the positional relationship between the target hand and the second target object. In other embodiments, the first target object may be other objects depending on the application scenario. For example, if the image to be detected is a road scene, both the first and second target objects may be vehicles.

[0032] It should be noted that the target hand refers to the hand of a specific object. For example, if the image to be processed captures a scene of user A's activity, then the target hand is user A's hand. When two hands are detected in the image, if both hands belong to user A, both hands can be considered as target hands, and then each target hand is checked to determine whether it is holding an object. When more than two hands are detected in the image, the hand of the target object is identified and designated as the target hand, while the other hands are not considered as detection objects.

[0033] In this embodiment, the second target object refers to the object held by the target hand. For this solution, the objects that the user's target hand might hold in the application scenario are predicted in advance. For example, the second target object could be a mobile phone, a cup, a tablet, a remote control, a book, etc. The image to be processed is then detected to obtain the position of the second target object in the image, that is, the second region occupied by the second target object in the image.

[0034] For example, in one embodiment, before obtaining the first region where the first target object is located and the second region where the second target object is located in the image to be processed, the method further includes: detecting the image to be processed according to the object recognition model to determine whether the first target object and the second target object exist in the image to be processed; when the first target object and the second target object exist in the image to be processed, performing the operation of obtaining the first region where the first target object is located and the second region where the second target object is located in the image to be processed.

[0035] This object detection model is trained on a pre-built convolutional neural network. Multiple images containing a target hand and a second target object are used as sample images, and the categories of the target hand and the second target object are labeled in the sample images as tag data. For all sample images, some may contain only the target hand, some only the second target object, and others may contain both. The pre-built convolutional neural network is trained using these labeled sample images to determine the weight parameters. This convolutional neural network with determined weight parameters is then used as the object detection model to detect whether a first target object and a second target object exist in the image to be processed. For example, in the application scenario of this embodiment, the object the user might be holding is a mobile phone; therefore, the mobile phone is considered the second target object. Thus, when training this object detection model, images containing various types of mobile phones can be used as some sample images.

[0036] If the image to be processed is determined to contain both a target hand and a second target object, subsequent position detection operations continue. Conversely, if the image to be processed contains only a target hand and no second target object, the detection ends, and no further detection operations are required. In the case of gesture recognition, gesture recognition operations can continue to be performed on the image to be processed. If the image to be processed contains no target hand, the detection ends, and no further operations are required.

[0037] In another embodiment, the object detection model can also be used to determine the positions of the first and second object in the image to be processed, given the presence of both. During the training phase of this object detection model, the sample images contain both category and location labels.

[0038] After acquiring the image to be processed, it is input into the object detection model for detection and processing, and the model outputs the locations of the target hand and the second target object in the image. For example, the object detection model can output the locations of the target hand and the second target object in the form of bounding boxes, and then determine the region where the bounding box corresponding to the target hand is located as the first region where the target hand is located, and the region where the bounding box corresponding to the second target object is located as the second region where the second target object is located.

[0039] For example, in another embodiment, the image to be processed is input into a target detection model for detection, and a first candidate region where the first target object is located and a second candidate region where the second target object is located are output. This includes: inputting the image to be processed into a first target detection model for detection processing, and outputting a first region where the target hand in the image to be processed is located; inputting the image to be processed into a second target detection model for detection processing, and outputting a second region where the second target object in the image to be processed is located.

[0040] Please see Figure 2 , Figure 2 This is a schematic diagram of a target detection scene in the image processing method provided in this application embodiment. In this embodiment, the target detection model includes a first target detection model and a second target detection model. That is, for the target hand and the second target object, a corresponding target detection model is trained respectively, and is used to detect the first region where the target hand is located and the second region where the second target object is located. The specific detection principle is the same as in the previous embodiment, and will not be repeated here.

[0041] 102. Calculate the intersection-union ratio of the first region and the second region.

[0042] After obtaining the first region occupied by the target hand and the second region occupied by the second target object in the image to be processed, the intersection-union ratio (IUGR) of the first and second regions in the image is calculated. Specifically, if the user's target hand is holding the second target object, then there will definitely be an overlap between the target hand and the second target object on the plane of the image to be processed, and the proportion of this overlap in their intersection will be greater than a certain threshold. Please refer to [link / reference]. Figure 3 , Figure 3 This is a schematic diagram illustrating a scenario where a mobile phone is held in hand. Figure 3 It can be seen that part of the phone is obscured by the phone, meaning that the area where the phone is located and the area where the target's hand is located partially overlap.

[0043] The threshold can be determined by analyzing actual images of the object being held. This threshold will be referred to as the first preset threshold in the following text.

[0044] In one embodiment, calculating the intersection-union ratio of the first region and the second region includes: determining the intersecting region and the merged region of the first region and the second region; obtaining a first number of pixels in the intersecting region and a second number of pixels in the merged region; calculating the ratio between the first number and the second number, and using the ratio as the intersection-union ratio between the first region and the second region.

[0045] In this embodiment, the intersection-union ratio (IUGR) of the first and second regions is calculated by counting the number of pixels. Specifically, the intersecting region between the first and second regions is obtained, and the first number of pixels in the intersecting region is counted. Then, the union of the first and second regions, i.e., the merged region, is obtained, and the second number of pixels in the merged region is counted. Finally, the ratio between the first and second numbers is calculated, and this ratio is used as the IUGR between the first and second regions.

[0046] Alternatively, in another embodiment, calculating the intersection-merger ratio of the first region and the second region includes: determining the intersecting region and the merged region of the first region and the second region; calculating the first area of ​​the first region and calculating the second area of ​​the merged region; calculating the ratio of the first area to the second area, and using the ratio as the intersection-merger ratio of the first region and the second region.

[0047] In this embodiment, the intersection-union ratio (IUR) between the first and second regions is calculated using their length and width dimensions. When the target detection model outputs the region where the target hand is located, it can represent the target bounding box position using the vertex + length and width dimensions. Then, based on the vertex + length and width dimensions of the first and second regions, the vertex positions and length and width dimensions of the intersecting region can be determined. The first area of ​​the intersecting region is then calculated based on its length and width dimensions, and the second area of ​​the merged region is calculated based on the length and width dimensions of the first and second regions. The ratio of the first area to the second area is then used as the IUR between the first and second regions.

[0048] 103. When the crossover ratio is greater than the first preset threshold, calculate the difference in occupancy between the first region and the second region in the image to be processed.

[0049] When the cross-union ratio (CUR) is greater than a first preset threshold, it can be preliminarily determined that there may be overlap between the first target object and the second target object. For example, the target hand may be holding the second target object. Assuming the first preset threshold is 30%, if the calculated CUR is 40%, it can be preliminarily determined that the target hand in the image to be processed may be holding the second target object. If the calculated CUR is 5%, it can be determined that the target hand in the image to be processed is holding the second target object.

[0050] However, since the scene being captured is a three-dimensional space, while the image reflects two-dimensional information, when the target hand and the second target object are both located in the direction perpendicular to the image plane but do not intersect, they may still appear to have a relatively large overlap on the image plane. Therefore, to avoid misjudging this spatial overlap as a holding phenomenon, when the overlap ratio is detected to be greater than a first preset threshold, the difference in occupancy rates between the first and second regions in the image to be processed is calculated, and further judgment is made based on this difference. Here, occupancy rate refers to the proportion of the object in the captured image.

[0051] For example, in one embodiment, calculating the difference in the proportion of the first region and the second region in the image to be processed includes: calculating a first area proportion of the first region in the image to be processed, and calculating a second area proportion of the second region in the image to be processed; calculating the absolute value of the difference between the first area proportion and the second area proportion, and using the absolute value of the difference as the difference in the proportion of the first region and the second region in the image to be processed.

[0052] In this embodiment, the proportions of the target hand and the second target object in the two-dimensional image to be processed are calculated. For the target hand, the proportion in the image to be processed represents the occupancy rate of the target hand in the space captured by the image. A first area proportion of the first region in the entire image is calculated, a second area proportion of the second region in the entire image is calculated, and the absolute value of the difference between the first and second area proportions is calculated. This absolute value is used as the difference between the proportions of the first and second regions in the image to be processed.

[0053] 104. Generate the positional relationship information between the first target object and the second target object based on the difference.

[0054] After determining the difference in the proportions of the first and second regions in the image to be processed, the spatial relationship between the first and second target objects can be determined based on this difference. Taking a mobile phone as an example, when holding the phone, the difference between the proportion of the phone in the image to be processed and the proportion of the hand in the image should be within a reasonable range; that is, the absolute value of the difference should not exceed a specific threshold. In practical applications, a reasonable threshold can be determined by analyzing actual images of the object being held. This threshold will be referred to as the second preset threshold below.

[0055] Conversely, when the target hand and the second target are both located in the line of sight perpendicular to the image plane but do not intersect, the difference in their proportions in the image to be processed will be relatively large. For example, if the target hand is closer to the lens, its first area proportion is larger, while the second target is farther from the lens, its second area proportion is smaller, and the final calculated difference will be greater than the second preset threshold.

[0056] Based on the principle described above, if the calculated difference is greater than or equal to the second preset threshold, it can be determined that the target hand in the image to be processed is not holding the second target object. Conversely, if the difference is less than the second preset threshold, it can be determined that the target hand in the image to be processed is holding the second target object.

[0057] In one embodiment, a preset threshold corresponding to a second target is determined based on the mapping relationship between multiple preset thresholds and multiple preset target objects, and the preset threshold corresponding to the second target object is determined as the second preset threshold.

[0058] Since different second target objects may have different sizes—for example, a tablet computer's volume and area are significantly larger than a mobile phone's—multiple preset thresholds can be pre-set for different second target objects in practical applications. After identifying the second target object, the corresponding preset threshold is selected from these multiple preset thresholds and determined as the second preset threshold. This method allows for flexible adjustment of the second preset threshold to adapt to the specific scenario and improve the accuracy of image processing.

[0059] Alternatively, in another embodiment, the first area proportion of the first region in the image to be processed is calculated; based on the mapping relationship between multiple preset thresholds and multiple preset areas, a preset threshold corresponding to the first area proportion is determined, and the preset threshold corresponding to the first area proportion is determined as the second preset threshold.

[0060] In scenarios involving holding objects, the difference in occupancy can be significant when the target hand (or a secondary target object) occupies different proportions in the image. Using a fixed second preset threshold for determining whether an object is being held may result in relatively low accuracy. For example, if both the target hand and the secondary target object are close to the lens, and the hand is holding an object, they will occupy a large proportion of the image, and the difference in proportion may also be large. In this case, a relatively large preset threshold can be chosen as the second preset threshold. Conversely, if both the target hand and the secondary target object are far from the lens, and the hand is holding an object, they will occupy a small proportion of the image, and the difference in proportion may also be small. In this case, a relatively small preset threshold can be chosen as the second preset threshold to accurately determine whether an object is being held. Based on this principle, multiple preset thresholds are pre-defined with mapping relationships between multiple first area proportions (or second area proportions). The corresponding preset threshold is determined as the second preset threshold based on the first area proportion calculated for the current scene and the mapping relationship, thereby improving the accuracy of image processing.

[0061] Understandably, after calculating the intersection-union ratio (IUR) of the first and second regions, the method further includes: when the IUR is less than or equal to a first preset threshold, determining that the target hand in the image to be processed is not holding the second target object. After calculating the difference in the proportions of the first and second regions in the image to be processed, the method further includes: when the difference is greater than or equal to a second preset threshold, determining that the target hand in the image to be processed is not holding the second target object; or, when the difference is greater than or equal to a second preset threshold, determining that the target hand and the second target object in the image to be processed are independent of each other.

[0062] In practice, this application is not limited by the execution order of the described steps. Without causing conflicts, some steps may be performed in other orders or simultaneously.

[0063] As can be seen from the above, the image processing method provided in this application embodiment obtains the regions where the first target object and the second target object are located in the image to be processed, calculates the intersection-union ratio (IUGR) between the two regions, and when the IUGR is greater than a first preset threshold, it can be determined that the first target object and the second target object are on the same line of sight. In this case, the difference in the proportion of the two regions in the image to be processed is calculated, and the positional relationship information between the first target object and the second target object is generated based on the difference. The scheme of this application embodiment, based on two dimensions—the IUGR of the two regions and the difference in the proportion of the two regions in the image—improves the accuracy of detecting the positional relationship between objects.

[0064] Furthermore, when the real-time solution of this application is applied to object holding detection scenarios, the region where the target hand is located and the region where the target object is located are identified from the image to be detected. The overlap rate between the two regions is then calculated. When the overlap rate is greater than a first preset threshold, it can be determined that the target hand and the target object are on the same line of sight. In this case, the difference in occupancy of the two regions in the image to be detected is calculated. When this difference is less than a second preset threshold, it can be determined that the two are on the same horizontal space, rather than overlapping front and back in space. This allows for the determination that the target hand is holding the target object, achieving accurate detection of the target hand holding the object based on the image. In addition, based on accurate object holding recognition, compared with conventional recognition methods, this solution does not require extracting depth information of the hand and the target object to determine their absolute positions in space, thus saving software and hardware costs.

[0065] In one embodiment, obtaining the first region where the first target object is located and the second region where the second target object is located in the image to be processed includes: inputting the image to be processed into a target detection model for detection, and outputting the first candidate region where the first target object is located and the second candidate region where the second target object is located; identifying the first contour of the first target object in the first candidate region, and generating the first region corresponding to the first target object based on the first contour; identifying the second contour of the second target object in the second candidate region, and generating the second region corresponding to the second target object based on the second contour.

[0066] In this embodiment, in order to improve the accuracy of subsequent cross-union ratio calculation, more accurate regional information of the target hand and target object is obtained from the image to be processed by recognizing the contours of the target hand and target object.

[0067] Specifically, after acquiring the image to be processed, it is first input into a pre-trained object detection model. This model calculates and outputs a first candidate region containing the target hand and a second candidate region containing the target object, either as a bounding box or coordinates. Then, contour recognition is performed on the image of the first candidate region to obtain the first contour of the target hand. For example, the image of the first candidate region is converted to a grayscale image, and edge detection is performed to obtain image edge information. The first contour of the target hand is then extracted from the image edge information. Finally, the closed region formed by this first contour is used as the first region corresponding to the target hand. It can be understood that if the generated first contour has interrupted parts, the interrupted parts can be supplemented according to the direction of the lines at both ends of the interrupted part to form a closed region.

[0068] The solution in this embodiment can more accurately identify the areas occupied by the target hand and the target object in the image to be processed, improve the accuracy of the cross-union ratio, and thus improve the accuracy of the positional relationship between the target hand and the second target object.

[0069] In one embodiment, an image processing apparatus is also provided. See also... Figure 4 , Figure 4 This is a schematic diagram of the structure of an image processing apparatus 300 provided in an embodiment of this application. The image processing apparatus 300 includes:

[0070] The target detection module 301 is used to acquire the first region where the first target object is located and the second region where the second target object is located in the image to be processed;

[0071] The first calculation module 302 is used to calculate the intersection-union ratio of the first region and the second region;

[0072] The second calculation module 303 is used to calculate the difference in the proportion of the first region and the second region in the image to be processed when the cross-union ratio is greater than the first preset threshold.

[0073] Image processing module 304 is used to generate positional relationship information between the first target object and the second target object based on the difference.

[0074] In some embodiments, the target detection module 301 is further configured to detect the image to be processed according to the object recognition model to determine whether there is a first target object and a second target object in the image to be processed; when there is a first target object and a second target object in the image to be processed, the operation of obtaining the first region where the first target object is located and the second region where the second target object is located in the image to be processed is performed.

[0075] In some embodiments, the target detection module 301 is further configured to input the image to be processed into the target detection model for detection, output a first candidate region where the first target is located and a second candidate region where the second target is located; identify a first contour of the first target in the first candidate region and generate a first region corresponding to the first target based on the first contour; and identify a second contour of the second target in the second candidate region and generate a second region corresponding to the second target based on the second contour.

[0076] In some embodiments, the first calculation module 302 is used to determine the intersection region and the merged region of the first region and the second region; obtain a first number of pixels in the intersection region and a second number of pixels in the merged region; calculate the ratio between the first number and the second number, and use the ratio as the intersection-merge ratio of the first region and the second region.

[0077] In some embodiments, the first calculation module 302 is used to determine the intersecting region and the merged region between the first region and the second region; calculate the first area of ​​the first region and the second area of ​​the merged region; calculate the ratio of the first area to the second area, and use the ratio as the intersection-merger ratio of the first region and the second region.

[0078] In some embodiments, the second calculation module 303 is used to calculate the first area ratio of the first region in the image to be processed, and to calculate the second area ratio of the second region in the image to be processed.

[0079] Calculate the absolute value of the difference between the first area proportion and the second area proportion, and use the absolute value of the difference as the difference between the proportions of the first region and the second region in the image to be processed.

[0080] In some embodiments, the first target object is a target hand; the image processing module 304 is used to determine that the target hand is holding the second target object when the difference is less than a second preset threshold.

[0081] In some embodiments, the device further includes:

[0082] The threshold determination module is used to determine the preset threshold corresponding to the second target object based on the mapping relationship between multiple preset thresholds and multiple preset target objects, and to determine the preset threshold corresponding to the second target object as the second preset threshold.

[0083] In some embodiments, the device further includes:

[0084] The threshold determination module is used to calculate the first area proportion of the first region in the image to be processed; determine the preset threshold corresponding to the first area proportion according to the mapping relationship between multiple preset thresholds and multiple preset areas, and determine the preset threshold corresponding to the first area proportion as the second preset threshold.

[0085] In some embodiments, the first target object is a target hand; the target detection module 301 is further configured to determine the target hand from among the multiple hands when there are multiple hands in the image to be processed.

[0086] It should be noted that the image processing apparatus provided in this application embodiment and the image processing method in the above embodiment belong to the same concept. The image processing apparatus can implement any of the methods provided in the image processing method embodiment. For details of its implementation process, please refer to the image processing method embodiment, which will not be repeated here.

[0087] As can be seen from the above, the image processing apparatus proposed in this application acquires the regions where the first target object and the second target object are located in the image to be processed, calculates the intersection-union ratio (IUGR) between the two regions, and determines that the first target object and the second target object are on the same line of sight when the IUGR is greater than a first preset threshold. In this case, the difference in the proportion of the two regions in the image to be processed is calculated, and positional relationship information between the first target object and the second target object is generated based on the difference. The scheme of this application improves the accuracy of detecting the positional relationship between objects by performing detection based on two dimensions: the IUGR of the two regions and the difference in the proportion of the two regions in the image.

[0088] This application also provides an electronic device, which can be a terminal, such as a smartphone, tablet computer, laptop computer, touch screen, game console, personal computer (PC), personal digital assistant (PDA), or other terminal device. Please refer to... Figure 5 , Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device 400 includes a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, and a computer program stored in the memory 402 and executable on the processor. The processor 401 and the memory 402 are electrically connected. Those skilled in the art will understand that the electronic device structure shown in the figure does not constitute a limitation on the electronic device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0089] The processor 401 is the control center of the electronic device 400. It connects various parts of the electronic device 400 through various interfaces and lines. By running or loading software programs and / or modules stored in the memory 402, and calling data stored in the memory 402, it performs various functions of the electronic device 400 and processes data, thereby monitoring the electronic device 400 as a whole.

[0090] In this embodiment, the processor 401 in the electronic device 400 loads the instructions corresponding to the processes of one or more applications into the memory 402 according to the following steps, and the processor 401 runs the applications stored in the memory 402 to realize various functions:

[0091] Obtain the first region where the first target object is located and the second region where the second target object is located in the image to be processed;

[0092] Calculate the intersection-union ratio of the first region and the second region;

[0093] When the crossover ratio is greater than the first preset threshold, the difference between the proportions of the first region and the second region in the image to be processed is calculated.

[0094] The positional relationship information between the first target and the second target is generated based on the difference.

[0095] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0096] Optional, such as Figure 5 As shown, the electronic device 400 also includes: a touch display screen 403, a radio frequency circuit 404, an audio circuit 405, an input unit 406, and a power supply 407. The processor 401 is electrically connected to the touch display screen 403, the radio frequency circuit 404, the audio circuit 405, the input unit 406, and the power supply 407. Those skilled in the art will understand that... Figure 5 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0097] The touch display screen 403 can be used to display a graphical user interface (GUI) and receive operation commands generated by the user interacting with the GUI. The touch display screen 403 may include a display panel and a touch panel. The display panel can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the electronic device. These graphical user interfaces can be composed of graphics, text, icons, video, and any combination thereof. The touch panel can be used to collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch panel), and generate corresponding operation commands, which then execute the corresponding program. Optionally, the touch panel may include a touch detection device and a touch controller. The touch detection device detects the user's touch location and the signal generated by the touch operation, transmitting the signal to the touch controller. The touch controller receives touch information from the touch detection device, converts it into touch point coordinates, sends it to the processor 401, and can receive and execute commands from the processor 401.

[0098] The radio frequency circuit 404 can be used to transmit and receive radio frequency signals to establish wireless communication with network devices or other electronic devices, and to transmit and receive signals with network devices or other electronic devices.

[0099] Audio circuit 405 can be used to provide an audio interface between a user and an electronic device via a speaker and a microphone. Audio circuit 405 can convert received audio data into electrical signals and transmit them to the speaker, where the speaker converts them into sound signals for output. Conversely, the microphone converts collected sound signals into electrical signals, which are then received by audio circuit 405, converted back into audio data, and then processed by processor 401 before being transmitted via radio frequency circuit 404 to, for example, another electronic device, or output to memory 402 for further processing. Audio circuit 405 may also include an earphone jack to provide communication between peripheral headphones and electronic devices.

[0100] The input unit 406 can be used to receive input numbers, characters, or user characteristic information (such as fingerprints, iris, facial information, etc.), and to generate keyboard, mouse, joystick, optical, or trackball signal inputs related to user settings and function control.

[0101] Power supply 407 is used to supply power to various components of electronic device 400. Optionally, power supply 407 can be logically connected to processor 401 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. Power supply 407 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0102] although Figure 5 As not shown in the diagram, the electronic device 400 may also include a camera, sensor, wireless fidelity module, Bluetooth module, etc., which will not be described in detail here.

[0103] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0104] As can be seen from the above, the electronic device provided in this embodiment acquires the regions where the first target object and the second target object are located in the image to be processed, calculates the intersection-union ratio (IUGR) between the two regions, and when the IUGR is greater than a first preset threshold, it can be determined that the first target object and the second target object are on the same line of sight. In this case, the difference in the proportion of the two regions in the image to be processed is calculated, and the positional relationship information between the first target object and the second target object is generated based on the difference. The scheme of this application embodiment, based on two dimensions—the IUGR of the two regions and the difference in the proportion of the two regions in the image—improves the accuracy of detecting the positional relationship between objects.

[0105] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0106] Therefore, embodiments of this application provide a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the steps of any of the image processing methods provided in embodiments of this application. For example, the computer program can perform the following steps:

[0107] Obtain the first region where the first target object is located and the second region where the second target object is located in the image to be processed;

[0108] Calculate the intersection-union ratio of the first region and the second region;

[0109] When the crossover ratio is greater than the first preset threshold, the difference between the proportions of the first region and the second region in the image to be processed is calculated.

[0110] The positional relationship information between the first target and the second target is generated based on the difference.

[0111] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0112] The storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0113] Since the computer program stored in the storage medium can execute the steps of any of the image processing methods provided in the embodiments of this application, the beneficial effects that any of the image processing methods provided in the embodiments of this application can achieve can be realized. For details, please refer to the previous embodiments, which will not be repeated here.

[0114] The foregoing has provided a detailed description of an image processing method, apparatus, device, and computer-readable storage medium provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. An image processing method, characterized in that, include: Obtain the first region where the first target object is located and the second region where the second target object is located in the image to be processed; Calculate the intersection-union ratio of the first region and the second region; When the cross-union ratio is greater than a first preset threshold, the difference between the proportions of the first region and the second region in the image to be processed is calculated. The positional relationship information between the first target and the second target is generated based on the difference. The first target object is the target hand; The step of generating positional relationship information between the first target and the second target based on the difference includes: When the difference is less than a second preset threshold, it is determined that the target hand is holding the second target object; Before determining that the target hand is holding the second target object when the difference is less than a second preset threshold, the method further includes: Based on the mapping relationship between multiple preset thresholds and multiple preset target objects, a preset threshold corresponding to the second target object is determined, and the preset threshold corresponding to the second target object is determined as the second preset threshold; or... Calculate the first area percentage of the first region in the image to be processed; determine the preset threshold corresponding to the first area percentage based on the mapping relationship between multiple preset thresholds and multiple preset areas, and determine the preset threshold corresponding to the first area percentage as the second preset threshold.

2. The method as described in claim 1, characterized in that, Before obtaining the first region where the first target object is located and the second region where the second target object is located in the image to be processed, the method further includes: The image to be processed is detected according to the object recognition model to determine whether a first target object and a second target object exist in the image to be processed. When a first target object and a second target object exist in the image to be processed, the operation of obtaining the first region where the first target object is located and the second region where the second target object is located in the image to be processed is performed.

3. The method as described in claim 1, characterized in that, The step of obtaining the first region where the first target object is located and the second region where the second target object is located in the image to be processed includes: The image to be processed is input into the target detection model for detection, and the first candidate region where the first target object is located and the second candidate region where the second target object is located are output. Identify the first contour of the first target object in the first candidate region, and generate a first region corresponding to the first target object based on the first contour; Identify the second contour of the second target object in the second candidate region, and generate the second region corresponding to the second target object based on the second contour.

4. The method as described in claim 1, characterized in that, The calculation of the intersection-union ratio of the first region and the second region includes: Determine the intersecting and merging regions of the first and second regions; Obtain a first number of pixels in the intersecting region and a second number of pixels in the merged region; Calculate the ratio between the first quantity and the second quantity, and use the ratio as the intersection-union ratio of the first region and the second region.

5. The method as described in claim 1, characterized in that, The calculation of the intersection-union ratio of the first region and the second region includes: Determine the intersecting and merging regions of the first and second regions; Calculate the first area of ​​the first region, and calculate the second area of ​​the merged region; Calculate the ratio of the first area to the second area, and use the ratio as the intersection-union ratio of the first region and the second region.

6. The method as described in claim 1, characterized in that, The calculation of the difference in the proportion of the first region and the second region in the image to be processed includes: Calculate the first area percentage of the first region in the image to be processed, and calculate the second area percentage of the second region in the image to be processed; Calculate the absolute value of the difference between the first area proportion and the second area proportion, and use the absolute value of the difference as the difference between the proportions of the first region and the second region in the image to be processed.

7. The method as described in claim 2, characterized in that, The step of detecting the image to be processed according to the object recognition model to determine whether a first target object and a second target object exist in the image to be processed includes: When there are multiple hands in the image to be processed, the target hand is determined from the multiple hands.

8. An image processing apparatus, characterized in that, include: The target detection module is used to obtain the first region where the first target object is located and the second region where the second target object is located in the image to be processed; The first calculation module is used to calculate the intersection-union ratio of the first region and the second region; The second calculation module is used to calculate the difference between the proportions of the first region and the second region in the image to be processed when the cross-union ratio is greater than the first preset threshold. The image processing module is used to generate positional relationship information between the first target and the second target based on the difference; The first target object is a hand, and the image processing module is used for: When the difference is less than a second preset threshold, it is determined that the target hand is holding the second target object; The image processing device further includes a threshold determination module; The threshold determination module is used to determine the preset threshold corresponding to the second target object based on the mapping relationship between multiple preset thresholds and multiple preset target objects, and to determine the preset threshold corresponding to the second target object as the second preset threshold. or, The threshold determination module is used to calculate the first area percentage of the first region in the image to be processed. Based on the mapping relationship between multiple preset thresholds and multiple preset areas, a preset threshold corresponding to the first area percentage is determined, and the preset threshold corresponding to the first area percentage is determined as the second preset threshold.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the image processing method as described in any one of claims 1 to 7.

10. An electronic device, characterized in that, The electronic device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the image processing method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Target object detection method and device and electronic equipment

    CN113705643A

  • Object detection method and object detection device

    WO2019235050A1