Image processing method, apparatus, device, storage medium, and program
By using image processing technology to automatically detect user interface images, the problem of low efficiency caused by text exceeding the bounding box is solved, and efficient detection of objects exceeding the bounding box is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- DOUYIN VISION CO LTD
- Filing Date
- 2022-04-20
- Publication Date
- 2026-05-08
AI Technical Summary
Existing technologies for detecting text overflow in user interfaces are inefficient and require significant manpower and time.
Image processing technology is used to process the user interface image through electronic devices, the detection area is determined by sliding the detection box, and the area of the object exceeding the box is automatically detected based on the detection results.
It enables automatic detection of objects that exceed the bounding box in the user interface, improving detection efficiency and saving manpower and time costs.
Smart Images

Figure CN116958180B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to an image processing method, apparatus, device, storage medium, and program. Background Technology
[0002] With technological advancements, human-computer interaction is becoming increasingly common. The interface that users directly interact with during interactions with applications on electronic devices is typically called the User Interface (UI).
[0003] The user interface includes various interface elements, such as views, windows, dialog boxes, menus, buttons, and labels. Some interface elements can display text. In some cases, text may exceed the boundaries of the interface element. This can affect the aesthetics of the interface and may also obscure other interface elements.
[0004] Currently, the main method is to manually inspect the user interface for text exceeding the bounding box, which is inefficient and requires significant manpower and time. Summary of the Invention
[0005] This disclosure provides an image processing method, apparatus, device, storage medium, and program to improve the detection efficiency of hyperframed objects, thereby reducing labor and time costs.
[0006] In a first aspect, embodiments of this disclosure provide an image processing method, including:
[0007] Obtain the interface image corresponding to the first user interface, wherein the interface image includes at least one display object;
[0008] At least one detection box is slid across the interface image to determine N detection regions in the interface image, and the detection results of the N detection regions are determined respectively; the detection result of each detection region includes: the coordinate information of the detection region, a first probability that the detection region is the region occupied by the superframe object and / or a second probability that the detection region is the region occupied by the superframe portion of the superframe object; N is an integer greater than 1;
[0009] Based on the detection results of the N detection areas, a first target area occupied by the superframe object and / or a second target area occupied by the superframe portion of the superframe object are determined in the interface image; wherein, the superframe object is a display object in the interface image whose at least part of the area exceeds the preset display boundary, and the superframe portion is the part of the superframe object that exceeds the preset display boundary.
[0010] In a second aspect, embodiments of this disclosure provide an image processing apparatus, comprising:
[0011] An acquisition module is used to acquire an interface image corresponding to the first user interface, wherein the interface image includes at least one display object;
[0012] A detection module is used to slide at least one detection box on the interface image to determine N detection regions in the interface image, and to determine the detection results of the N detection regions respectively; the detection result of each detection region includes: the coordinate information of the detection region, a first probability that the detection region is the area occupied by the superframe object and / or a second probability that the detection region is the area occupied by the superframe portion of the superframe object; where N is an integer greater than 1;
[0013] The determining module is used to determine, based on the detection results of the N detection regions, a first target region occupied by the superframe object and / or a second target region occupied by the superframe portion of the superframe object in the interface image; wherein, the superframe object is a display object in the interface image whose at least part of the region exceeds the preset display boundary, and the superframe portion is the part of the superframe object that exceeds the preset display boundary.
[0014] Thirdly, embodiments of this disclosure provide an electronic device, including: a processor and a memory;
[0015] The memory stores computer-executed instructions;
[0016] The processor executes computer execution instructions stored in the memory, causing the processor to perform the image processing method as described in the first aspect above.
[0017] Fourthly, embodiments of this disclosure provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the image processing method described in the first aspect above.
[0018] Fifthly, embodiments of this disclosure provide a computer program product, including a computer program that, when executed by a processor, implements the image processing method described in the first aspect above.
[0019] The image processing method, apparatus, device, storage medium, and program provided in this disclosure include: acquiring an interface image corresponding to a first user interface, the interface image including at least one display object; sliding at least one detection box across the interface image to determine N detection regions in the interface image, and determining the detection results for each of the N detection regions; the detection result for each detection region includes: coordinate information of the detection region, a first probability that the detection region is an area occupied by a superframe object, and / or a second probability that the detection region is an area occupied by the superframe portion of the superframe object; and determining, based on the detection results of the N detection regions, a first target area occupied by the superframe object, and / or a second target area occupied by the superframe portion of the superframe object, in the interface image. In the above process, by performing detection processing on the interface image corresponding to the first user interface, automatic detection of superframe objects in the first user interface is achieved, improving detection efficiency and saving labor and time costs. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 A schematic diagram of a user interface provided for an embodiment of this disclosure;
[0022] Figure 2 This is a schematic diagram illustrating a process for detecting hyperframe objects in a user interface, provided by an embodiment of the present disclosure.
[0023] Figure 3 This is a schematic flowchart of an image processing method provided in an embodiment of the present disclosure;
[0024] Figure 4 This is a schematic diagram illustrating the detection processing of an interface image using a detection box, as provided in an embodiment of the present disclosure.
[0025] Figure 5 A schematic diagram of a system architecture for detecting hyperframe objects in a user interface, provided by an embodiment of this disclosure;
[0026] Figure 6 This is a schematic diagram of the structure of a preset model provided in an embodiment of the present disclosure;
[0027] Figure 7 This is a schematic diagram of another preset model provided in an embodiment of the present disclosure;
[0028] Figure 8A schematic flowchart of another image processing method provided in an embodiment of this disclosure;
[0029] Figure 9 A schematic diagram of an image processing procedure provided in an embodiment of this disclosure;
[0030] Figure 10 This is a schematic diagram of the structure of an image processing apparatus provided in an embodiment of the present disclosure;
[0031] Figure 11 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0032] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0033] The technical solution provided in this disclosure can be used to detect hyperframe objects in a user interface (UI). The user interface refers to the interface that a user directly interacts with during the interaction between the user and an application in an electronic device. A user interface may include one or more interface elements. Interface elements are a series of elements in the user interface that meet the user's interaction needs, including but not limited to: views, windows, dialog boxes, menus, buttons, and labels.
[0034] In some scenarios, applications may be launched and released in multiple countries around the world. When an application is launched and released in different countries, it is necessary to translate the text in the user interface into the corresponding language. Due to the differences in the languages of different countries, the translated text in the user interface may exceed the text frame. This will not only affect the aesthetics of the interface, but may also obscure other interface elements.
[0035] To facilitate understanding, the following will be combined with... Figure 1 Let's illustrate with examples. Figure 1 This is a schematic diagram of a user interface provided in an embodiment of this disclosure. Taking a game interface as an example, as shown... Figure 1 As shown, the user interface includes eight interface elements, numbered 101 to 108. Interface elements 101, 105, 106, and 107 are labels used to statically display text. Interface elements 102, 103, 104, and 108 are buttons that can be clicked by the user.
[0036] See also Figure 1 Text can be displayed / carried on interface elements of the user interface. For example, the text displayed on interface element 101 is "game item", the text displayed on interface element 102 is "item A", and the text displayed on interface element 105 is "item A has characteristics A1 and A2", etc.
[0037] In some cases, text displayed on UI elements may extend beyond the boundaries of those elements. For example, Figure 1 In the example, the text displayed on interface element 106 reads "Item B has characteristic B1 and is used in situation B2," but this text extends beyond the boundaries of interface element 106. For another example, Figure 1 In the middle, the text displayed on interface element 108 is "Click to learn more about props", and this text goes beyond the boundaries of interface element 108.
[0038] In this embodiment of the disclosure, the phenomenon of "text exceeding the boundary of interface elements" is referred to as the "text exceeding the boundary" phenomenon. Text exhibiting this phenomenon is called "over-boundary text." For example, Figure 1 In the text, "Item B has the characteristics of B1 and can be used in the case of B2" can be called "over-frame text", and "Click to learn more about items" can also be called "over-frame text".
[0039] It should be understood that, in addition to displaying / carrying text, interface elements of a user interface can also display / carry other content, such as images, icons, and other interface elements. In this embodiment of the disclosure, the content displayed / carrying on interface elements is collectively referred to as display objects. For each type of display object, there may be a situation where it exceeds the preset display boundary (e.g., the boundary of the interface element displaying / carrying the display object). In this embodiment of the disclosure, the phenomenon of "display object exceeding the preset display boundary" is referred to as "object exceeding the frame". Object exceeding the frame can include one or more of the following: text exceeding the frame, image exceeding the frame, icon exceeding the frame, interface element exceeding the frame, etc.
[0040] Furthermore, display objects that exhibit the phenomenon of "exceeding the preset display boundary" are called "over-frame objects." In other words, an over-frame object refers to a display object where at least a portion extends beyond the preset display boundary. The portion of an over-frame object that extends beyond the preset display boundary is called the "over-frame portion." For example, Figure 1 In the superframe object "Item B has the characteristic of B1, and is used in the case of B2", the three characters "used below" extend beyond the boundary of the interface element 106. Therefore, the text "used below" within this superframe object can be considered the superframe portion. For example, Figure 1In the text "Click to learn more about props", the two characters "props" in the super-framed object exceed the boundary of the interface element 108. Therefore, the text "props" in the super-framed object can be called the super-framed part.
[0041] Currently, when inspecting hyperframe objects in the user interface, manual visual inspection is required, which is inefficient and incurs significant manpower and time costs.
[0042] In this embodiment of the disclosure, image processing technology can be used to process the image corresponding to the user interface by an electronic device. Figure 2 This is a schematic diagram illustrating a process for detecting hyperframe objects in a user interface, as provided in an embodiment of this disclosure. Figure 2 As shown, a screenshot or photograph of the user interface to be detected is taken, and the interface image is input into an electronic device. The electronic device processes the interface image and marks the area occupied by the hyperframed object in the interface image (e.g., Figure 2 (in dashed boxes 201 and 203), and / or, the area occupied by the super-boundary portion of the super-boundary object (e.g., Figure 2 (dashed bounding boxes 202 and 204 in the text). This enables automatic detection of hyperframe objects in the user interface, improving detection efficiency and saving labor and time costs.
[0043] The technical solutions provided in this disclosure will be described in detail below with reference to several specific embodiments. The specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0044] Figure 3 This is a schematic flowchart illustrating an image processing method provided in an embodiment of this disclosure. The method of this embodiment can be derived from... Figure 2 The electronic devices within it perform the execution. For example... Figure 3 As shown, the method in this embodiment includes:
[0045] S301: Obtain the interface image corresponding to the first user interface, wherein the interface image includes at least one display object.
[0046] In this embodiment, the first user interface is the user interface to be detected, meaning that this embodiment needs to detect whether there are any objects exceeding the bounding box in the first user interface. The interface image includes one or more display objects, and each display object can be displayed (or contained, carried, or presented) within a certain interface element.
[0047] In this embodiment of the disclosure, the displayed object can be one or more of the following: text, image, icon, interface element, etc. It should be noted that when the displayed object is text, it can refer to a line of text (i.e., a text line) or a paragraph of text (i.e., a text segment). A line of text may include one or more characters. For example, Figure 1 In this context, the text displayed by each interface element is treated as a display object.
[0048] The interface image corresponding to the first user interface can be a screenshot or a photograph of the first user interface. The content displayed in the interface image is the same as the content displayed in the first user interface.
[0049] S302: At least one detection box is slid across the interface image to determine N detection regions in the interface image, and the detection results of the N detection regions are determined respectively; the detection result of each detection region includes: the coordinate information of the detection region, the first probability that the detection region is the region occupied by the superframe object and / or the second probability that the detection region is the region occupied by the superframe part of the superframe object; N is an integer greater than 1.
[0050] To facilitate understanding, the following will be combined with... Figure 4 Let's illustrate with examples. Figure 4 This is a schematic diagram illustrating the detection processing of an interface image using a detection box, as provided in an embodiment of this disclosure. Figure 4 As shown, assume the interface image comprises 10 pixels in width and 10 pixels in height, meaning the interface image size is 10*10. Assume the detection box has a width of 5 and a height of 3, meaning the detection box size is 5*3. The detection box slides across the interface image; during the slide, at each position, the detection box defines a detection region on the interface image. In this way, multiple detection regions can be generated. For example, Figure 4 In the diagram, the black rectangle is the detection box. As the detection box moves, the area enclosed by the detection box (i.e., the shaded area) is a detection region.
[0051] For example, the pixel (i, j) in the interface image is used as the center of the detection box, and the area covered / enclosed by the detection box on the interface image forms a detection region. Here, i takes values of 1, 2, ..., 10, and j takes values of 1, 2, ..., 10. This results in a total of 100 detection regions.
[0052] For each detection region, detection can be performed to determine its category. In this disclosure, there are two categories of detection regions: one is the "region occupied by the hyperframe object," and the other is the "region occupied by the hyperframe portion of the hyperframe object." Thus, the detection result for each detection region can include: the coordinate information of the detection region, and a first probability that the detection region is the "region occupied by the hyperframe object." Alternatively, the detection result for each detection region can include: the coordinate information of the detection region, and a second probability that the detection region is the "region occupied by the hyperframe portion of the hyperframe object." Or, the detection result for each detection region can include: the coordinate information of the detection region, a first probability that the detection region is the "region occupied by the hyperframe object," and a second probability that the detection region is the "region occupied by the hyperframe portion of the hyperframe object."
[0053] The coordinate information of the detection area can be represented in various ways, and this embodiment does not limit this. For example, it can be represented by the coordinates of the center point (x, y), width w, and height h of the detection area; it can also be represented by the coordinates of the top left corner vertex (x1, y1) and the bottom right corner vertex (x2, y2) of the detection area; or it can be represented by the offset (Δx, Δy), width w, and height h of the center point coordinates of the detection area relative to a certain preset coordinate.
[0054] S303: Based on the detection results of the N detection areas, determine the first target area occupied by the super-frame object and / or the second target area occupied by the super-frame portion of the super-frame object in the interface image; wherein, the super-frame object is a display object in the interface image whose at least part of the area exceeds the preset display boundary, and the super-frame portion is the part of the super-frame object that exceeds the preset display boundary.
[0055] In other words, based on the detection results of the N detection regions, a first target region occupied by the hyperframe object is determined in the interface image. Alternatively, based on the detection results of the N detection regions, a second target region occupied by the hyperframe portion of the hyperframe object is determined in the interface image. Or, based on the detection results of the N detection regions, both the first target region occupied by the hyperframe object and the second target region occupied by the hyperframe portion of the hyperframe object are determined in the interface image.
[0056] The following explanation uses the simultaneous determination of the first target region and the second target region as an example.
[0057] In this embodiment, the area occupied by the hyperframe object is referred to as the first target area, for example... Figure 2The dashed boxes 201 and 203 are used in the definition. The area occupied by the super-frame portion of the super-frame object is called the second target area, such as dashed boxes 202 and 204. It should be understood that in actual applications, there may be multiple super-frame objects in the first user interface. In this case, there may be multiple first target areas and multiple second target areas determined in S303.
[0058] It should be understood that since the detection results of N detection areas have been determined, the first target area occupied by the hyperframe object and the second target area occupied by the hyperframe portion of the hyperframe object can be determined from the N detection areas based on the detection results of the N detection areas.
[0059] For example, the N detection regions can be filtered based on the first probability and the second probability corresponding to the N detection regions. For instance, the detection regions with lower probabilities can be excluded to obtain the first target region and the second target region.
[0060] It should be noted that the above Figure 4 The example shown uses a single detection box for illustration. In some possible implementations, there can be multiple detection boxes of different sizes. This allows for the detection of multiple regions of different sizes, thus enabling the detection of hyperframe objects of varying sizes within the interface image and improving the accuracy of the detection results.
[0061] When multiple detection boxes of different sizes exist, the N detection regions detected in S302 will contain detection regions of different sizes. Thus, in S303, when determining the first target region and the second target region from the N detection regions, in addition to filtering the N detection regions based on the first probability and the second probability corresponding to each detection region, the N detection regions can also be filtered based on the coordinate information of each detection region.
[0062] When filtering N detection regions based on their coordinate information, the filtering can be based on the overlap between different detection regions. For example, if detection region 1 completely covers detection region 2, then detection region 2 can be deleted. Alternatively, if the intersection-to-union (IoU) ratio between detection regions 3 and 4 is greater than or equal to a preset threshold, then either detection region 3 or detection region 4 can be deleted.
[0063] Optionally, the N detection regions can be first filtered based on the first and second probabilities corresponding to each detection region; then, the remaining detection regions can be filtered based on the coordinate information of each detection region. Alternatively, the N detection regions can be filtered first based on the coordinate information of each detection region; then, the remaining detection regions can be filtered based on the first and second probabilities corresponding to each detection region. Alternatively, the N detection regions can be filtered simultaneously based on both the first and second probabilities corresponding to each detection region and the coordinate information of each detection region.
[0064] Among some possible implementations, the first target region and the second target region can be determined from N detection regions in the following way:
[0065] (1) Based on the first probability and the second probability corresponding to the N detection regions, a plurality of first candidate regions and a plurality of second candidate regions are determined in the N detection regions; wherein, the first probability corresponding to the first candidate region is greater than the first threshold, and the second probability corresponding to the second candidate region is greater than the second threshold.
[0066] In other words, for each detection region, if the first probability corresponding to that detection region is greater than a first threshold, then that detection region is determined as a first candidate region. If the second probability corresponding to that detection region is greater than a second threshold, then that detection region is determined as a second candidate region. In this way, the first target region will be determined from each of the first candidate regions, and the second target region will be determined from each of the second candidate regions.
[0067] (2) Based on the coordinate information of the plurality of first candidate regions and the coordinate information of the plurality of second candidate regions, the first target region is determined in the plurality of first candidate regions, and the second target region is determined in the plurality of second candidate regions.
[0068] For example, based on the coordinate information of multiple first candidate regions, overlapping first candidate regions can be eliminated, and then the first target region can be determined from the remaining first candidate regions. Similarly, based on the coordinate information of multiple second candidate regions, overlapping second candidate regions can be eliminated, and then the second target region can be determined from the remaining second candidate regions.
[0069] In this implementation, it is equivalent to performing two rounds of filtering on N detection regions to determine the first target region and the second target region. The first filtering (i.e., step (1)) is based on the first probability and the second probability, which can filter out some detection regions with lower probabilities. The second filtering (i.e., step (2)) is based on coordinate information, which can filter out some detection regions with overlapping relationships. Thus, the first target region and the second target region are finally determined, ensuring the accuracy of the detection results.
[0070] The image processing method provided in this embodiment includes: acquiring an interface image corresponding to a first user interface, the interface image including at least one display object; sliding at least one detection box on the interface image to determine N detection regions in the interface image, and determining the detection results of each of the N detection regions, the detection result of each detection region including: the coordinate information of the detection region, a first probability that the detection region is the area occupied by a hyperframe object, and / or a second probability that the detection region is the area occupied by the hyperframe portion of the hyperframe object; and determining a first target area occupied by the hyperframe object and / or a second target area occupied by the hyperframe portion of the hyperframe object in the interface image based on the detection results of the N detection regions. In the above process, by performing detection processing on the interface image corresponding to the first user interface, automatic detection of hyperframe objects in the first user interface is achieved, improving detection efficiency and saving labor and time costs.
[0071] Figure 3 In the illustrated embodiment, S302 and S303 can also be implemented using a preset model. That is, the preset model uses at least one detection box to slide across the interface image to determine N detection regions in the interface image, and determines the detection results for each of the N detection regions. Then, based on the detection results of the N detection regions, the preset model determines a first target region occupied by the hyperframe object and / or a second target region occupied by the hyperframe portion of the hyperframe object in the interface image.
[0072] The preset model can be trained using machine learning methods. For example, the preset model is trained on multiple sets of training samples, each set of training samples including: a sample image, the region occupied by a hyperframe object in the sample image, and the region occupied by the hyperframe portion of the hyperframe object.
[0073] Figure 5 This is a schematic diagram of a system architecture for detecting hyperframe objects in a user interface, provided as an embodiment of this disclosure. Figure 5 As shown, the system architecture may include a training device and an execution device. The training device can train a pre-defined model using multiple sets of training samples from a database. This pre-defined model can then be deployed to the execution device.
[0074] When it is necessary to detect whether there is a bounding object in the first user interface, the interface image corresponding to the first user interface can be input into the execution device. The execution device performs detection processing on the interface image using a preset model, and obtains the detection results of N detection regions. Then, based on the detection results of the N detection regions, the execution device determines the first target region occupied by the bounding object and / or the second target region occupied by the bounding portion of the bounding object in the interface image using the preset model.
[0075] In some embodiments, the execution device and the training device can be the same device. In other embodiments, the execution device and the training device can be different electronic devices.
[0076] In this disclosure, a preset model is used to detect hyperframe objects in the user interface; the preset model can also be called a hyperframe object detector.
[0077] The following is combined with Figure 6 and Figure 7 The structure of the preset model and the process by which the preset model processes interface images are explained in detail.
[0078] Figure 6 This is a schematic diagram of a preset model provided in an embodiment of this disclosure. Figure 6 As shown, the preset model may include: a feature extraction network, a feature fusion network, and a hyperframe detection network. The number of feature fusion networks and hyperframe detection networks can be multiple.
[0079] pass Figure 6 The process of processing the interface image using the preset model shown is as follows:
[0080] (1) Extract features from the interface image at K different scales to obtain feature maps corresponding to each of the K scales. K is an integer greater than 1.
[0081] For example, see Figure 6 Assuming K=3, a feature extraction network is used to extract features from the interface image at three different scales, resulting in feature maps corresponding to the first, second, and third scales. The first scale is larger than the second, and the second scale is larger than the third. For example, the first scale is 52*52, the second scale is 26*26, and the third scale is 13*13.
[0082] Furthermore, for each scale of feature map, at least one detection box is used to perform detection processing on that feature map. Taking the i-th scale as an example, at least one detection box is used to slide across the feature map corresponding to the i-th scale to determine N in the feature map corresponding to the i-th scale. iEach detection area is defined, and the N values are determined accordingly. i The detection results of each detection area; the N i It is an integer greater than 1. See steps (2) to (4) below for details.
[0083] (2) For the Kth scale, at least one detection box is slid across the feature map corresponding to the Kth scale to determine N in the feature map corresponding to the Kth scale. k Each detection area is defined, and the N values are determined accordingly. k The test results for each testing area.
[0084] For example, see Figure 6 Assume K = 3. The feature map corresponding to the 3rd scale is processed using the hyperframe detection network 3 to obtain the detection results for N3 detection regions.
[0085] (3) For the i-th scale, the feature map corresponding to the i-th scale and the feature map corresponding to the (i+1)-th scale are fused to obtain the fused feature map corresponding to the i-th scale. At least one detection box is then slid across the fused feature map corresponding to the i-th scale to determine N in the fused feature map corresponding to the i-th scale. i Each detection area is defined, and the N values are determined accordingly. i The test results for each testing area.
[0086] In step (3) above, i takes the values K-1, K-2, ..., 2, 1 in sequence.
[0087] For example, see Figure 6 The feature map corresponding to the second scale and the feature map corresponding to the third scale are fused using feature fusion network 2 to obtain the fused feature map corresponding to the second scale. Further, the fused feature map corresponding to the second scale is processed by hyperframe detection network 2 to obtain detection results for N2 detection regions.
[0088] See also Figure 6 The feature map corresponding to the first scale and the feature map corresponding to the second scale (it should be noted that either the original feature map corresponding to the second scale or the fused feature map corresponding to the second scale can be used) are fused through feature fusion network 1 to obtain the fused feature map corresponding to the first scale. Further, the fused feature map corresponding to the first scale is processed by hyperframe detection network 1 to obtain the detection results of N1 detection regions.
[0089] In this embodiment, the feature map corresponding to the i-th scale and the feature map corresponding to the (i+1)-th scale are fused to obtain the fused feature map corresponding to the i-th scale, making the feature map corresponding to the i-th scale more accurate.
[0090] Optionally, the feature map corresponding to the i-th scale and the feature map corresponding to the (i+1)-th scale can be fused to obtain the fused feature map corresponding to the i-th scale. This can be done as follows: the feature map corresponding to the (i+1)-th scale is upsampled to obtain a sampled feature map, which has the same size as the feature map corresponding to the i-th scale. The sampled feature map and the feature map corresponding to the i-th scale are then fused to obtain the fused feature map corresponding to the i-th scale.
[0091] Thus, based on Figure 6 The preset model shown outputs detection results for N = N3 + N2 + N1 detection areas.
[0092] In the above process, for the i-th scale, the number N of the detection regions output by the hyperframe detection network i is... i Related to the i-th scale. The following section combines... Figure 6 Examples are given for the three scales in the text.
[0093] Assuming the third scale is 13*13, when the hyperframe detection network 3 performs detection processing, it uses one detection box, taking each point in the feature map corresponding to the third scale as the center point of the detection box. In this way, the hyperframe detection network 3 can obtain detection results for N3 = 13*13 regions. If M detection boxes are used, the hyperframe detection network 3 can obtain detection results for N3 = 13*13*M regions.
[0094] Assuming the second scale is 26*26, when the super-bounded detection network 2 performs detection processing, it uses one detection box and takes each point in the feature map corresponding to the fused second scale as the center point of the detection box. In this way, the super-bounded detection network 2 can obtain detection results for N² = 26*26 detection regions. If M detection boxes are used, the super-bounded detection network 2 can obtain detection results for N² = 26*26*M detection regions.
[0095] Assuming the first scale is 52*52, when the super-bounded detection network 1 performs detection processing, it uses one detection box and takes each point in the feature map corresponding to the fused first scale as the center point of the detection box. In this way, the super-bounded detection network 1 can obtain detection results for N1 = 52*52 detection regions. If M detection boxes are used, the super-bounded detection network 1 can obtain detection results for N1 = 52*52*M detection regions.
[0096] Figure 7 A schematic diagram of another preset model provided in an embodiment of this disclosure. The following is in conjunction with... Figure 7 Let's illustrate with examples.
[0097] See Figure 7Assuming the interface image includes 416 pixels in the width direction, 416 pixels in the height direction, and each pixel includes 3 channels in the channel direction, then the interface image can be denoted as (416, 416, 3).
[0098] The interface image is input into a feature extraction network, which includes convolutional unit 1 and residual units 1 through 5. See also... Figure 7 :
[0099] The interface image (416, 416, 3) is input into convolutional unit 1. Convolutional unit 1 performs channel expansion processing on the interface image to obtain image features (416, 416, 32). For example, convolutional unit 1 can be a Conv2D 32×3×3.
[0100] Residual unit 1 performs downsampling and channel expansion processing on the image feature (416, 416, 32) to obtain the image feature (208, 208, 64). For example, residual unit 1 can be a 1×64 Residual Block.
[0101] Residual unit 2 performs downsampling and channel expansion processing on the image features (208, 208, 64) to obtain image features (104, 104, 128). For example, residual unit 2 can be a Residual Block 2×128.
[0102] Residual unit 3 performs downsampling and channel expansion processing on the image features (104, 104, 128) to obtain image features (52, 52, 256). For example, residual unit 3 can be a Residual Block 8×256.
[0103] Residual unit 4 performs downsampling and channel expansion processing on the image features (52, 52, 256) to obtain image features (26, 26, 512). For example, residual unit 4 can be a Residual Block 8×512.
[0104] Residual unit 5 performs downsampling and channel expansion processing on image features (26, 26, 512) to obtain image features (13, 13, 1024). For example, residual unit 5 can be a Residual Block 4×1024.
[0105] In this embodiment, taking K=3 as an example, three feature scales are used. Assume the first scale is 52*52, the second scale is 26*26, and the third scale is 13*13. That is, after passing through the feature extraction network, the feature map corresponding to the first scale (52, 52, 256), the feature map corresponding to the second scale (26, 26, 512), and the feature map corresponding to the third scale (13, 13, 1024) are extracted.
[0106] Furthermore, the feature map (13, 13, 1024) corresponding to the third scale is input into the hyperframe detection network 3. The hyperframe detection network 3 includes convolutional unit 2 and detection unit 1. After processing by convolutional unit 2 and detection unit 1, the hyperframe detection network 3 outputs the detection result (1313, 21). For example, convolutional unit 2 can be Conv2D Block 5L 1024, and detection unit 1 can be Conv2D 3×3+Conv2D 1×1.
[0107] The feature maps (26, 26, 512) and (13, 13, 1024) corresponding to the second scale are input into the feature fusion network 2. The feature fusion network 2 includes an upsampling unit 1 and a fusion unit 1. The upsampling unit 1 upsamples the feature map (13, 13, 1024) corresponding to the third scale to obtain the sampled features (26, 26, 256) at the second scale. The fusion unit 1 fuses the feature map (26, 26, 512) corresponding to the second scale with the sampled features (26, 26, 256) output by the upsampling unit 1 to obtain the fused feature map (26, 26, 768) corresponding to the second scale. For example, the upsampling unit 1 can be Conv2D+UpSampling2D, and the fusion unit 1 can be Concat.
[0108] The fused feature map (26, 26, 768) corresponding to the second scale is input into the hyperframe detection network 2. The hyperframe detection network 2 includes a convolutional unit 3 and a detection unit 2. The convolutional unit 3 performs channel reduction processing on the fused feature map (26, 26, 768) corresponding to the second scale to obtain the feature map (26, 26, 256) of the second scale. The detection unit 2 performs hyperframe detection processing on the feature map (26, 26, 256) output by the convolutional unit 3 to obtain the detection result (26, 26, 21). For example, the convolutional unit 3 can be a Conv2D Block 5L 256, and the detection unit 2 can be a Conv2D 3×3+Conv2D 1×1.
[0109] The feature map (52, 52, 256) corresponding to the first scale and the feature map (26, 26, 256) corresponding to the second scale output by convolutional unit 3 are input into feature fusion network 1. Feature fusion network 1 includes upsampling unit 2 and fusion unit 2. Upsampling unit 2 upsamples the feature map (26, 26, 256) corresponding to the second scale output by convolutional unit 3 to obtain the sampled features (52, 52, 128) for the first scale. Fusion unit 2 fuses the feature map (52, 52, 256) corresponding to the first scale and the sampled features (52, 52, 128) output by upsampling unit 2 to obtain the fused feature map (52, 52, 384) corresponding to the first scale. For example, upsampling unit 2 can be Conv2D+UpSampling2D, and fusion unit 2 can be Concat.
[0110] The fused feature map (52, 52, 384) corresponding to the first scale is input into the hyperframe detection network 1. The hyperframe detection network 1 includes a convolutional unit 4 and a detection unit 3. The convolutional unit 4 performs channel reduction processing on the fused feature map (52, 52, 384) corresponding to the first scale to obtain the feature map (52, 52, 128) of the first scale. The detection unit 3 performs hyperframe detection processing on the feature map (52, 52, 128) output by the convolutional unit 4 to obtain the detection result (52, 52, 21). For example, the convolutional unit 4 can be a Conv2D Block 5L 128, and the detection unit 3 can be a Conv2D 3×3+Cony2D 1×1.
[0111] After the above processing, the preset model outputs detection results at three scales: (52, 52, 21), (26, 26, 21), and (1313, 21). In these three results, 21 = 3 * (4 + 2 + 1). Here, 3 represents the number of detection boxes. 4 represents the coordinate information of the detection region (e.g., center point coordinates (x, y), width w, and height h). 2 represents the first probability that the detection region is the "region occupied by the bounding box object," the second probability that the detection region is the "region occupied by the bounding box portion of the bounding box object," and 1 indicates whether the detection region is a bounding box object.
[0112] In other words, based on Figure 7 The preset model shown uses 3 detection boxes at each scale, and outputs the detection results of 52*52*3+26*26*3+13*13*3 detection regions. The detection result of each detection region is (4+2+1)=7 dimensions.
[0113] In this embodiment, feature extraction processing is performed on the interface image at multiple scales, and then super-boundary detection is performed based on the image features at each scale. In this way, super-boundary objects at different scales can be detected, improving the accuracy of the detection results.
[0114] Figure 8 This is a schematic flowchart illustrating another image processing method provided in an embodiment of this disclosure. Figure 8 As shown, the method in this embodiment includes:
[0115] S801: Obtain the interface image corresponding to the first user interface, wherein the interface image includes at least one display object.
[0116] S802: At least one detection box is slid across the interface image to determine N detection regions in the interface image, and the detection results of the N detection regions are determined respectively; the detection result of each detection region includes: the coordinate information of the detection region, the first probability that the detection region is the area occupied by the superframe object, and the second probability that the detection region is the area occupied by the superframe part of the superframe object; N is an integer greater than 1.
[0117] It should be understood that the specific implementation of S801 and S802 in this embodiment is similar to that in the previous embodiments, and will not be described in detail here.
[0118] S803: Identify the display objects in the interface image to obtain the coordinate information of S display objects.
[0119] In this embodiment, the display object recognition algorithm corresponding to the type of the hyperframe object to be detected can be used to identify the display objects in the interface image and obtain the coordinate information of S display objects.
[0120] For example, assuming we need to detect superframe text, we would perform text recognition on the interface image to obtain the coordinate information of each text object in the interface image. For instance, we could use an Optical Character Recognition (OCR) algorithm to perform text recognition on the interface image to obtain the coordinate information of each text line in the interface image.
[0121] For example, if it is necessary to detect super-frame interface elements, the interface image can be used to identify interface elements and obtain the coordinate information of each interface element in the interface image.
[0122] It should be noted that this embodiment does not limit the execution order of S802 and S803. The execution order of the two can be interchanged, or the two can be executed simultaneously.
[0123] S804: Based on the first probability and the second probability corresponding to the N detection regions, determine a plurality of first candidate regions and a plurality of second candidate regions among the N detection regions; wherein, the first probability corresponding to the first candidate region is greater than a first threshold, and the second probability corresponding to the second candidate region is greater than a second threshold.
[0124] In other words, for each detection region, if the first probability corresponding to that detection region is greater than a first threshold, then that detection region is determined as a first candidate region. If the second probability corresponding to that detection region is greater than a second threshold, then that detection region is determined as a second candidate region. In this way, the first target region will be determined from each of the first candidate regions, and the second target region will be determined from each of the second candidate regions.
[0125] S805: Based on the coordinate information of the plurality of first candidate regions, the coordinate information of the plurality of second candidate regions, and the coordinate information of the S display objects, determine the first target region occupied by the super-frame object in the plurality of first candidate regions, and determine the second target region occupied by the super-frame portion of the super-frame object in the plurality of second candidate regions.
[0126] For example, the following principles can be used to filter multiple first candidate regions to obtain one or more first target regions.
[0127] (A1) If the intersection-union ratio (IUU) between two first candidate regions is greater than or equal to a preset threshold, it indicates that the two first candidate regions are overlapping regions, or mostly overlapping regions, and one of the two first candidate regions can be deleted. For example, the first candidate region with the lower probability can be deleted from the two first candidate regions.
[0128] (A2) If for a certain first candidate region, there is no overlap between the first candidate region and the S display objects identified in S803, then the first candidate region is deleted.
[0129] (A3) If for a certain first candidate region, there is no overlap between the first candidate region and all the second candidate regions, then the first candidate region is deleted.
[0130] (A4) If a first candidate region overlaps with a second candidate region, but the overlapping region is not located at the edge of the first candidate region, then the first candidate region is deleted.
[0131] For example, the following principles can be used to filter multiple second candidate regions to obtain one or more second target regions.
[0132] (B1) If the intersection-union ratio between two second candidate regions is greater than or equal to a preset threshold, it indicates that the two second candidate regions are overlapping regions, or mostly overlapping regions, and one of the two second candidate regions can be deleted. For example, the second candidate region with the lower probability can be deleted from the two second candidate regions.
[0133] (B2) If for a certain second candidate region, there is no overlap between the second candidate region and the S display objects identified in S803, then the second candidate region is deleted.
[0134] (B3) If a second candidate region does not overlap with any of the first candidate regions, then the second candidate region is deleted.
[0135] (B4) If a second candidate region overlaps with a first candidate region, but the overlapping region is not located at the edge of the first candidate region, then the second candidate region is deleted.
[0136] It should be noted that the filtering order of principles A1, A2, A3, and A4 above is not limited and can be in any order. Similarly, the filtering order of principles B1, B2, B3, and B4 above is not limited and can be in any order.
[0137] Thus, the first target region determined satisfies the following conditions:
[0138] (1) When the number of first target regions is greater than 1, the crossover ratio between any two first target regions is less than the preset threshold.
[0139] (2) The first target area has a first overlapping area with at least one of the S display objects.
[0140] (3) There is a second overlapping area between the first target area and the second target area, and the second overlapping area is located at the edge of the first target area.
[0141] Similarly, the identified second target region satisfies the following conditions:
[0142] (1) When the number of second target regions is greater than 1, the crossover ratio between any two second target regions is less than the preset threshold.
[0143] (2) The second target area has a first overlapping area with at least one of the S display objects.
[0144] (3) The second target area and the first target area have a second overlapping area, and the second overlapping area is located at the edge of the first target area.
[0145] exist Figure 8 Based on the embodiments shown, the following example illustrates the detection process of super-framed text. Figure 9 This is a schematic diagram illustrating an image processing procedure provided in an embodiment of this disclosure. Figure 9 As shown, suppose we need to detect super-frame text in the first user interface. The detection process is as follows:
[0146] (1) Take a screenshot or photograph of the first user interface to obtain the interface image.
[0147] (2) The interface image is processed by a preset model (or hyperframe text detector) to obtain the detection results of N detection areas.
[0148] It should be understood that the process of processing interface images can be found in [reference needed]. Figure 3 , Figure 6 , Figure 7 The detailed description of the illustrated embodiments is omitted here.
[0149] (3) By recognizing the text lines in the interface image based on the OCR algorithm, the coordinate information of S text lines is obtained.
[0150] (4) Based on the detection results of N detection areas and the coordinate information of S text lines, post-processing is performed to obtain the first target area where the super-frame text is located and the second target area occupied by the super-frame part of the super-frame text.
[0151] When performing post-processing, the following principles can be considered:
[0152] Principle 1: Both the first and second target regions should overlap with the region of one of the S text lines;
[0153] Principle 2: The second target area should overlap with a certain first target area, and the overlapping area should be located at the edge of the first target area, and the area of the overlapping area should be the same as or close to the area of the second target area.
[0154] It should be understood that the implementation method of post-processing in step (4) above can be found in [reference needed]. Figure 8 The detailed description of the illustrated embodiments is omitted here.
[0155] In this embodiment, by combining the detection results of the preset model (hyperframe text detector) with the text line detection results based on the OCR algorithm, the detection results of the preset model are corrected using the text line detection results of the OCR algorithm, thereby further improving the accuracy of the detection results.
[0156] The above embodiments detail the process of detecting hyperframe objects in a user interface using a preset model. This preset model needs to be pre-trained. This disclosure does not limit the training process of the preset model. It can be trained using multiple sets of training samples, each set including: a sample image, the region occupied by the hyperframe object in the sample image, and the region occupied by the hyperframe portion of the hyperframe object in the sample image.
[0157] In this embodiment of the disclosure, considering that the detection of hyperbounded objects is an anomaly detection, the training samples for training the preset model are limited. An algorithm can be used to automatically generate sample images containing hyperbounded objects. Taking hyperbounded text as an example, sample images can be generated in the following manner:
[0158] (1) Obtain the original image corresponding to the sample user interface.
[0159] (2) Use the OCR algorithm to identify the coordinate information of each text line in the original image.
[0160] (3) Use the interface element detection model to detect the original image and obtain the coordinate information of the interface elements in the original image.
[0161] (4) Randomly select interface elements carrying text lines in the original image and randomly select fonts from the font library.
[0162] (5) Erase the text lines within the selected interface elements.
[0163] (6) Based on the coordinate information of the interface element, use the selected font to write text information inside the interface element, and make the text information extend beyond the boundary of the interface element to obtain the sample image.
[0164] Through the above process, sample images containing bounding text are generated. Since the bounding text is automatically written using an algorithm, the area occupied by the bounding text in the sample image, as well as the area occupied by the bounding portion of the bounding text, can be determined. Furthermore, the sample image, the area occupied by the bounding text in the sample image, and the area occupied by the bounding portion of the bounding text can be used to train a model and obtain a pre-trained preset model.
[0165] In this embodiment, by using an algorithm to automatically generate training samples, the difficulty of collecting training samples is reduced, and the manpower and time costs required for manual annotation of training samples are reduced, thereby improving the training efficiency of the preset model.
[0166] Figure 10 This is a schematic diagram of the structure of an image processing apparatus provided in an embodiment of this disclosure. The image processing apparatus provided in this embodiment can be in the form of software and / or hardware. Figure 10As shown, the image processing apparatus 1000 provided in this embodiment may include: an acquisition module 1001, a detection module 1002, and a determination module 1003. Wherein,
[0167] The acquisition module 1001 is used to acquire an interface image corresponding to the first user interface, wherein the interface image includes at least one display object;
[0168] The detection module 1002 is used to slide at least one detection box on the interface image to determine N detection regions in the interface image, and to determine the detection results of the N detection regions respectively; the detection result of each detection region includes: the coordinate information of the detection region, a first probability that the detection region is the area occupied by the super-frame object and / or a second probability that the detection region is the area occupied by the super-frame portion of the super-frame object; N is an integer greater than 1;
[0169] The determining module 1003 is used to determine, based on the detection results of the N detection areas, a first target area occupied by the super-frame object and / or a second target area occupied by the super-frame portion of the super-frame object in the interface image; wherein, the super-frame object is a display object in the interface image whose at least part of the area exceeds the preset display boundary, and the super-frame portion is the part of the super-frame object that exceeds the preset display boundary.
[0170] In some possible implementations, the detection module 1002 is specifically used for:
[0171] The interface image is subjected to feature extraction at K different scales to obtain feature maps corresponding to each of the K scales; where K is an integer greater than 1.
[0172] At least one detection box is slid across the feature map corresponding to the i-th scale to determine N in the feature map corresponding to the i-th scale. i Each detection area is defined, and the N values are determined accordingly. i The detection results of each detection area; the N i It is an integer greater than 1;
[0173] Where i takes the values K, K-1, K-2, ..., 1 in sequence, and N = N K +N k-1 +…+N1.
[0174] In some possible implementations, when i < K, the detection module 1002 is specifically used for:
[0175] The feature map corresponding to the i-th scale and the feature map corresponding to the (i+1)-th scale are fused to obtain the fused feature map corresponding to the i-th scale.
[0176] At least one detection box is slid across the fused feature map corresponding to the i-th scale to determine N in the fused feature map corresponding to the i-th scale. i Each detection area is defined, and the N values are determined accordingly. i The test results for each testing area.
[0177] In some possible implementations, the i-th scale is larger than the (i+1)-th scale; the detection module 1002 is specifically used for:
[0178] Upsampling is performed on the feature map corresponding to the (i+1)th scale to obtain a sampled feature map, wherein the sampled feature map has the same size as the feature map corresponding to the i-th scale.
[0179] The sampled feature map and the feature map corresponding to the i-th scale are fused to obtain the fused feature map corresponding to the i-th scale.
[0180] In some possible implementations, the detection result for each detection region includes: the coordinate information of the detection region, a first probability that the detection region is the area occupied by the hyperframe object, and a second probability that the detection region is the area occupied by the hyperframe portion of the hyperframe object; the determination module 1003 is specifically used for:
[0181] Based on the first probability and the second probability corresponding to the N detection regions, a plurality of first candidate regions and a plurality of second candidate regions are determined in the N detection regions; wherein, the first probability corresponding to the first candidate region is greater than a first threshold, and the second probability corresponding to the second candidate region is greater than a second threshold;
[0182] Based on the coordinate information of the plurality of first candidate regions and the coordinate information of the plurality of second candidate regions, the first target region is determined in the plurality of first candidate regions, and the second target region is determined in the plurality of second candidate regions.
[0183] In some possible implementations, module 1003 is specifically used for:
[0184] The display objects in the interface image are identified to obtain the coordinate information of S display objects; where S is an integer greater than or equal to 1.
[0185] Based on the coordinate information of the plurality of first candidate regions, the coordinate information of the plurality of second candidate regions, and the coordinate information of the S display objects, the first target region is determined in the plurality of first candidate regions, and the second target region is determined in the plurality of second candidate regions.
[0186] In some possible implementations, the first target region satisfies the following two conditions:
[0187] The first target area overlaps with at least one of the S display objects in a first overlapping area;
[0188] The first target region and the second target region have a second overlapping region, and the second overlapping region is located at the edge of the first target region.
[0189] In some possible implementations, the second target region satisfies the following two conditions:
[0190] The second target area overlaps with at least one of the S display objects in a first overlapping area;
[0191] The second target region and the first target region have a second overlapping region, and the second overlapping region is located at the edge of the first target region.
[0192] In some possible implementations, the detection module 1002 is specifically used to: slide at least one detection box on the interface image using a preset model to determine N detection areas in the interface image, and determine the detection results of the N detection areas respectively;
[0193] The determining module 1003 is specifically used to: determine, based on the detection results of the N detection regions using the preset model, the first target region occupied by the hyperframe object in the interface image, and / or the second target region occupied by the hyperframe portion of the hyperframe object;
[0194] The preset model is obtained by training multiple sets of training samples. Each set of training samples includes: a sample image, the region occupied by the hyperframe object in the sample image, and the region occupied by the hyperframe portion of the hyperframe object in the sample image.
[0195] In some possible implementations, the display object can be any of the following: text, icon, image, or interface element.
[0196] The image processing apparatus provided in this embodiment can be used to execute the image processing method provided in any of the above method embodiments. Its implementation principle and technical effect are similar, and will not be described in detail here.
[0197] To implement the above embodiments, this disclosure also provides an electronic device.
[0198] refer to Figure 11The diagram illustrates a structural schematic of an electronic device 1100 suitable for implementing embodiments of the present disclosure. The electronic device 1100 can be a terminal device or a server. The terminal device can include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, personal digital assistants (PDAs), portable Android devices (PADs), portable media players (PMPs), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 11 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0199] like Figure 11 As shown, electronic device 1100 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 1101, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 1102 or a program loaded from storage device 1108 into random access memory (RAM) 1103. RAM 1103 also stores various programs and data required for the operation of electronic device 1100. The processing unit 1101, ROM 1102, and RAM 1103 are interconnected via bus 1104. Input / output (I / O) interface 1105 is also connected to bus 1104.
[0200] Typically, the following devices can be connected to I / O interface 1105: input devices 1106 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 1107 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1108 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1109. Communication device 1109 allows electronic device 1100 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 11 An electronic device 1100 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0201] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 1109, or installed from storage device 1108, or installed from ROM 1102. When the computer program is executed by processing device 1101, it performs the functions defined in the methods of embodiments of this disclosure.
[0202] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0203] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0204] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods shown in the above embodiments.
[0205] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0206] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0207] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".
[0208] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0209] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0210] In a first aspect, according to one or more embodiments of the present disclosure, an image processing method is provided, comprising:
[0211] Obtain the interface image corresponding to the first user interface, wherein the interface image includes at least one display object;
[0212] At least one detection box is slid across the interface image to determine N detection regions in the interface image, and the detection results of the N detection regions are determined respectively; the detection result of each detection region includes: the coordinate information of the detection region, a first probability that the detection region is the region occupied by the superframe object and / or a second probability that the detection region is the region occupied by the superframe portion of the superframe object; N is an integer greater than 1;
[0213] Based on the detection results of the N detection areas, a first target area occupied by the superframe object and / or a second target area occupied by the superframe portion of the superframe object are determined in the interface image; wherein, the superframe object is a display object in the interface image whose at least part of the area exceeds the preset display boundary, and the superframe portion is the part of the superframe object that exceeds the preset display boundary.
[0214] According to one or more embodiments of this disclosure, at least one detection box is slid across the interface image to determine N detection regions in the interface image, and the detection results of the N detection regions are determined respectively, including:
[0215] The interface image is subjected to feature extraction at K different scales to obtain feature maps corresponding to each of the K scales; where K is an integer greater than 1.
[0216] At least one detection box is slid across the feature map corresponding to the i-th scale to determine N in the feature map corresponding to the i-th scale. iEach detection area is defined, and the N values are determined accordingly. i The detection results of each detection area; the N i It is an integer greater than 1;
[0217] Where i takes the values K, K-1, K-2, ..., 1 in sequence, and N = N K +N k-1 +…+N1.
[0218] According to one or more embodiments of this disclosure, when i < K, at least one detection box is slid across the feature map corresponding to the i-th scale to determine N in the feature map corresponding to the i-th scale. i Each detection area is defined, and the N values are determined accordingly. i The detection results for each detection area include:
[0219] The feature map corresponding to the i-th scale and the feature map corresponding to the (i+1)-th scale are fused to obtain the fused feature map corresponding to the i-th scale.
[0220] At least one detection box is slid across the fused feature map corresponding to the i-th scale to determine N in the fused feature map corresponding to the i-th scale. i Each detection area is defined, and the N values are determined accordingly. i The test results for each testing area.
[0221] According to one or more embodiments of this disclosure, the i-th scale is larger than the (i+1)-th scale; the feature map corresponding to the i-th scale and the feature map corresponding to the (i+1)-th scale are fused to obtain the fused feature map corresponding to the i-th scale, including:
[0222] Upsampling is performed on the feature map corresponding to the (i+1)th scale to obtain a sampled feature map, wherein the sampled feature map has the same size as the feature map corresponding to the i-th scale.
[0223] The sampled feature map and the feature map corresponding to the i-th scale are fused to obtain the fused feature map corresponding to the i-th scale.
[0224] According to one or more embodiments of this disclosure, the detection result of each detection region includes: the coordinate information of the detection region, a first probability that the detection region is the region occupied by the hyperframe object, and a second probability that the detection region is the region occupied by the hyperframe portion of the hyperframe object.
[0225] Based on the detection results of the N detection regions, a first target region occupied by the hyperframe object and a second target region occupied by the hyperframe portion of the hyperframe object are determined in the interface image, including:
[0226] Based on the first probability and the second probability corresponding to the N detection regions, a plurality of first candidate regions and a plurality of second candidate regions are determined in the N detection regions; wherein, the first probability corresponding to the first candidate region is greater than a first threshold, and the second probability corresponding to the second candidate region is greater than a second threshold;
[0227] Based on the coordinate information of the plurality of first candidate regions and the coordinate information of the plurality of second candidate regions, the first target region is determined in the plurality of first candidate regions, and the second target region is determined in the plurality of second candidate regions.
[0228] According to one or more embodiments of this disclosure, determining the first target region from the plurality of first candidate regions and determining the second target region from the plurality of second candidate regions based on the coordinate information of the plurality of first candidate regions and the coordinate information of the plurality of second candidate regions includes:
[0229] The display objects in the interface image are identified to obtain the coordinate information of S display objects; where S is an integer greater than or equal to 1.
[0230] Based on the coordinate information of the plurality of first candidate regions, the coordinate information of the plurality of second candidate regions, and the coordinate information of the S display objects, the first target region is determined in the plurality of first candidate regions, and the second target region is determined in the plurality of second candidate regions.
[0231] According to one or more embodiments of this disclosure, the first target region satisfies the following two conditions:
[0232] The first target area overlaps with at least one of the S display objects in a first overlapping area;
[0233] The first target region and the second target region have a second overlapping region, and the second overlapping region is located at the edge of the first target region.
[0234] According to one or more embodiments of this disclosure, the second target region satisfies the following two conditions:
[0235] The second target area overlaps with at least one of the S display objects in a first overlapping area;
[0236] The second target region and the first target region have a second overlapping region, and the second overlapping region is located at the edge of the first target region.
[0237] According to one or more embodiments of this disclosure, at least one detection box is slid across the interface image to determine N detection regions in the interface image, and the detection results of the N detection regions are determined respectively, including:
[0238] By using a preset model, at least one detection box is slid across the interface image to determine N detection regions in the interface image, and the detection results of the N detection regions are determined respectively.
[0239] Based on the detection results of the N detection regions, a first target region occupied by the hyperframe object is determined in the interface image, and / or a second target region occupied by the hyperframe portion of the hyperframe object, including:
[0240] Based on the detection results of the N detection regions using the preset model, the first target region occupied by the hyperframe object and / or the second target region occupied by the hyperframe portion of the hyperframe object are determined in the interface image.
[0241] The preset model is obtained by training multiple sets of training samples. Each set of training samples includes: a sample image, the region occupied by the hyperframe object in the sample image, and the region occupied by the hyperframe portion of the hyperframe object in the sample image.
[0242] According to one or more embodiments of this disclosure, the display object is any one of the following: text, icon, image, interface element.
[0243] Secondly, according to one or more embodiments of the present disclosure, an image processing apparatus is provided, comprising:
[0244] An acquisition module is used to acquire an interface image corresponding to the first user interface, wherein the interface image includes at least one display object;
[0245] A detection module is used to slide at least one detection box on the interface image to determine N detection regions in the interface image, and to determine the detection results of the N detection regions respectively; the detection result of each detection region includes: the coordinate information of the detection region, a first probability that the detection region is the area occupied by the superframe object and / or a second probability that the detection region is the area occupied by the superframe portion of the superframe object; where N is an integer greater than 1;
[0246] The determining module is used to determine, based on the detection results of the N detection regions, a first target region occupied by the superframe object and / or a second target region occupied by the superframe portion of the superframe object in the interface image; wherein, the superframe object is a display object in the interface image whose at least part of the region exceeds the preset display boundary, and the superframe portion is the part of the superframe object that exceeds the preset display boundary.
[0247] According to one or more embodiments of this disclosure, the detection module is specifically used for:
[0248] The interface image is subjected to feature extraction at K different scales to obtain feature maps corresponding to each of the K scales; where K is an integer greater than 1.
[0249] At least one detection box is slid across the feature map corresponding to the i-th scale to determine N in the feature map corresponding to the i-th scale. i Each detection area is defined, and the N values are determined accordingly. i The detection results of each detection area; the N i It is an integer greater than 1;
[0250] Where i takes the values K, K-1, K-2, ..., 1 in sequence, and N = N K +N k-1 +…+N1.
[0251] According to one or more embodiments of this disclosure, when i < K, the detection module is specifically used for:
[0252] The feature map corresponding to the i-th scale and the feature map corresponding to the (i+1)-th scale are fused to obtain the fused feature map corresponding to the i-th scale.
[0253] At least one detection box is slid across the fused feature map corresponding to the i-th scale to determine N in the fused feature map corresponding to the i-th scale. i Each detection area is defined, and the N values are determined accordingly. i The test results for each testing area.
[0254] According to one or more embodiments of this disclosure, the i-th scale is larger than the (i+1)-th scale; the detection module is specifically used for:
[0255] Upsampling is performed on the feature map corresponding to the (i+1)th scale to obtain a sampled feature map, wherein the sampled feature map has the same size as the feature map corresponding to the i-th scale.
[0256] The sampled feature map and the feature map corresponding to the i-th scale are fused to obtain the fused feature map corresponding to the i-th scale.
[0257] According to one or more embodiments of this disclosure, the detection result of each detection region includes: coordinate information of the detection region, a first probability that the detection region is the region occupied by the hyperframe object, and a second probability that the detection region is the region occupied by the hyperframe portion of the hyperframe object; the determination module is specifically used for:
[0258] Based on the first probability and the second probability corresponding to the N detection regions, a plurality of first candidate regions and a plurality of second candidate regions are determined in the N detection regions; wherein, the first probability corresponding to the first candidate region is greater than a first threshold, and the second probability corresponding to the second candidate region is greater than a second threshold;
[0259] Based on the coordinate information of the plurality of first candidate regions and the coordinate information of the plurality of second candidate regions, the first target region is determined in the plurality of first candidate regions, and the second target region is determined in the plurality of second candidate regions.
[0260] According to one or more embodiments of this disclosure, the determining module is specifically used for:
[0261] The display objects in the interface image are identified to obtain the coordinate information of S display objects; where S is an integer greater than or equal to 1.
[0262] Based on the coordinate information of the plurality of first candidate regions, the coordinate information of the plurality of second candidate regions, and the coordinate information of the S display objects, the first target region is determined in the plurality of first candidate regions, and the second target region is determined in the plurality of second candidate regions.
[0263] According to one or more embodiments of this disclosure, the first target region satisfies the following two conditions:
[0264] The first target area overlaps with at least one of the S display objects in a first overlapping area;
[0265] The first target region and the second target region have a second overlapping region, and the second overlapping region is located at the edge of the first target region.
[0266] According to one or more embodiments of this disclosure, the second target region satisfies the following two conditions:
[0267] The second target area overlaps with at least one of the S display objects in a first overlapping area;
[0268] The second target region and the first target region have a second overlapping region, and the second overlapping region is located at the edge of the first target region.
[0269] According to one or more embodiments of this disclosure, the detection module is specifically used to: slide at least one detection box on the interface image using a preset model to determine N detection areas in the interface image, and determine the detection results of the N detection areas respectively;
[0270] The determination module is specifically used to: determine, based on the detection results of the N detection regions using the preset model, the first target region occupied by the hyperframe object in the interface image, and / or the second target region occupied by the hyperframe portion of the hyperframe object;
[0271] The preset model is obtained by training multiple sets of training samples. Each set of training samples includes: a sample image, the region occupied by the hyperframe object in the sample image, and the region occupied by the hyperframe portion of the hyperframe object in the sample image.
[0272] According to one or more embodiments of this disclosure, the display object is any one of the following: text, icon, image, interface element.
[0273] Thirdly, according to one or more embodiments of the present disclosure, an electronic device is provided, comprising: at least one processor and a memory;
[0274] The memory stores computer-executed instructions;
[0275] The at least one processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the image processing method as described in the first aspect and various possible designs of the first aspect.
[0276] Fourthly, according to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, wherein computer-executable instructions are stored therein, which, when executed by a processor, implement the image processing method described in the first aspect and various possible designs of the first aspect.
[0277] Fifthly, according to one or more embodiments of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the image processing method as described in the first aspect and various possible designs of the first aspect.
[0278] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0279] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0280] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. An image processing method, characterized in that, include: Obtain the interface image corresponding to the first user interface, wherein the interface image includes at least one display object; At least one detection box is slid across the interface image to determine N detection regions in the interface image, and the detection results of the N detection regions are determined respectively. The detection result for each detection region includes: the coordinate information of the detection region, the first probability that the detection region is the area occupied by the hyperbound object, and the second probability that the detection region is the area occupied by the hyperbound portion of the hyperbound object; where N is an integer greater than 1; Based on the first probability and the second probability corresponding to the N detection regions, a plurality of first candidate regions and a plurality of second candidate regions are determined in the N detection regions; the first probability corresponding to the first candidate region is greater than a first threshold, and the second probability corresponding to the second candidate region is greater than a second threshold; the display objects in the interface image are identified to obtain the coordinate information of S display objects; S is an integer greater than or equal to 1; based on the coordinate information of the plurality of first candidate regions, the coordinate information of the plurality of second candidate regions, and the coordinate information of the S display objects, a first target region occupied by a super-frame object is determined in the plurality of first candidate regions, and a second target region occupied by the super-frame portion of the super-frame object is determined in the plurality of second candidate regions; wherein, the super-frame object is a display object in the interface image whose at least part of its area exceeds a preset display boundary, and the super-frame portion is the part of the super-frame object that exceeds the preset display boundary; The second target area satisfies the following two conditions: the second target area has a first overlapping area with at least one of the S display objects; the second target area has a second overlapping area with the first target area, and the second overlapping area is located at the edge of the first target area.
2. The method according to claim 1, characterized in that, At least one detection bounding box is slid across the interface image to determine N detection regions in the interface image, and the detection results for each of the N detection regions are determined, including: The interface image is subjected to feature extraction at K different scales to obtain feature maps corresponding to each of the K scales; where K is an integer greater than 1. At least one detection box is slid across the feature map corresponding to the i-th scale to determine the feature map corresponding to the i-th scale. Each detection area is determined separately. The detection results of each detection area; It is an integer greater than 1; Where i takes the values K, K-1, K-2, ..., 1 in sequence. .
3. The method according to claim 2, characterized in that, When i < K, at least one detection box is used to slide on the feature map corresponding to the i-th scale, so as to determine detection regions in the feature map corresponding to the i-th scale, and respectively determine the detection results of the detection regions, including: The feature map corresponding to the i-th scale and the feature map corresponding to the (i+1)-th scale are fused to obtain the fused feature map corresponding to the i-th scale. At least one detection box is slid across the fused feature map corresponding to the i-th scale to determine the i-th scale feature map. Each detection area is determined separately. The test results for each testing area.
4. The method according to claim 3, characterized in that, The i-th scale is greater than the (i+1)-th scale; The feature map corresponding to the i-th scale and the feature map corresponding to the (i+1)-th scale are fused to obtain the fused feature map corresponding to the i-th scale, including: Upsampling is performed on the feature map corresponding to the (i+1)th scale to obtain a sampled feature map, wherein the sampled feature map has the same size as the feature map corresponding to the i-th scale. The sampled feature map and the feature map corresponding to the i-th scale are fused to obtain the fused feature map corresponding to the i-th scale.
5. The method according to claim 4, characterized in that, The first target region satisfies the following two conditions: The first target area overlaps with at least one of the S display objects in a first overlapping area; The first target region and the second target region have a second overlapping region, and the second overlapping region is located at the edge of the first target region.
6. The method according to any one of claims 1 to 5, characterized in that, At least one detection bounding box is slid across the interface image to determine N detection regions in the interface image, and the detection results for each of the N detection regions are determined, including: By using a preset model, at least one detection box is slid across the interface image to determine N detection regions in the interface image, and the detection results of the N detection regions are determined respectively. Based on the detection results of the N detection regions, a first target region occupied by the hyperframe object is determined in the interface image, and / or a second target region occupied by the hyperframe portion of the hyperframe object, including: Based on the detection results of the N detection regions using the preset model, the first target region occupied by the hyperframe object and / or the second target region occupied by the hyperframe portion of the hyperframe object are determined in the interface image. The preset model is obtained by training multiple sets of training samples. Each set of training samples includes: a sample image, the region occupied by the hyperframe object in the sample image, and the region occupied by the hyperframe portion of the hyperframe object in the sample image.
7. The method according to any one of claims 1 to 5, characterized in that, The display object can be any of the following: text, icon, image, or interface element.
8. An image processing apparatus, characterized in that, include: An acquisition module is used to acquire an interface image corresponding to the first user interface, wherein the interface image includes at least one display object; A detection module is used to slide at least one detection box on the interface image to determine N detection areas in the interface image, and to determine the detection results of the N detection areas respectively. The detection result for each detection region includes: the coordinate information of the detection region, the first probability that the detection region is the area occupied by the hyperbound object, and the second probability that the detection region is the area occupied by the hyperbound portion of the hyperbound object; where N is an integer greater than 1; The determining module is configured to determine multiple first candidate regions and multiple second candidate regions in the N detection regions based on the first probability and the second probability corresponding to the N detection regions; the first probability corresponding to the first candidate region is greater than a first threshold, and the second probability corresponding to the second candidate region is greater than a second threshold; identify display objects in the interface image to obtain coordinate information of S display objects; where S is an integer greater than or equal to 1; determine a first target region occupied by a superframe object in the multiple first candidate regions based on the coordinate information of the multiple first candidate regions, the coordinate information of the multiple second candidate regions, and the coordinate information of the S display objects, and determine a second target region occupied by the superframe portion of the superframe object in the multiple second candidate regions; wherein, the superframe object is a display object in the interface image whose at least part of its area exceeds a preset display boundary, and the superframe portion is the part of the superframe object that exceeds the preset display boundary; the second target region satisfies the following two conditions: the second target region has a first overlapping region with at least one of the S display objects; the second target region has a second overlapping region with the first target region, and the second overlapping region is located at the edge of the first target region.
9. An electronic device, characterized in that, include: Processor and memory; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the image processing method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, implement the image processing method as described in any one of claims 1 to 7.
11. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the image processing method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Image target detection method and device and storage medium
CN109815868A
Image detection method and device and computer readable storage medium
CN110807362A