Element positioning method and device and electronic equipment

By segmenting and mapping the target image, the mapping relationship between element regions and rows and columns is established, which solves the problem of inconsistent testing frameworks of different operating systems, and realizes the positioning ability to automatically adapt to changes in element attributes, reducing maintenance costs.

CN120013958APending Publication Date: 2025-05-16MICRO DREAM TECHTRONIC NETWORK TECH CHINACO
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411994820.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The testing frameworks of different operating systems are not unified, resulting in the need to maintain multiple sets of element positioning codes, which is high maintenance costs, and when element attributes change, the automatic positioning code is prone to failure to execute and needs to be updated frequently.

Method used

By obtaining the target image and its position index where the target element is located, image segmentation and element area mapping processing are performed, and the mapping relationship between element area and row and column binaries is established, and the target element is positioned based on this.

Benefits of technology

It realizes the ability to automatically adapt to property changes without modifying code positioning when the element rendering results remain unchanged, reduces maintenance costs, and is suitable for different operating systems, with only one operating system required to maintain.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013958A_ABST
    Figure CN120013958A_ABST
Patent Text Reader

Abstract

The invention discloses an element positioning method and device and electronic equipment, and relates to the field of element positioning. The element positioning method comprises the steps that a target image where a to-be-positioned target element is located and position indexes of the target element are obtained, the position indexes comprise row indexes and column indexes, and the position indexes of all the elements in the target image are defined based on the position layout of all the elements in the target image; carrying out image segmentation on the target image, carrying out element region mapping processing on each segmented region to obtain a position coordinate corresponding to each element region in the target image, and establishing a mapping relationship between the element region in the target image and the row and column two-tuple; and based on the position index of the target element and the mapping relationship between the element area in the target image and the row and column two-tuples, determining a position coordinate corresponding to the element area indicated by the position index of the target element so as to position the target element.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of element positioning, and in particular to an element positioning method, device and electronic equipment. Background Art

[0002] Different operating systems in the related technology use different test frameworks. Since the test frameworks are not universal, it is necessary to maintain multiple sets of element positioning codes for multiple test architectures, which has a high maintenance cost. In addition, since the framework locates the target element through the attributes of the element object when locating the interface element, the automatic positioning code will fail to execute when the element rendering effect remains unchanged but the attributes change, and maintenance and updates are forced, which has a high maintenance cost. Summary of the invention

[0003] The embodiments of the present application provide an element positioning method, device and electronic device to solve the problem in the related art that the test frameworks of different operating systems are not unified and multiple sets of element positioning codes need to be maintained.

[0004] In a first aspect, an embodiment of the present application provides an element positioning method, including: Acquire a target image where the target element to be located is located, and a position index of the target element, wherein the position index includes a row index and a column index, and the position index of each element in the target image is defined based on the position layout of each element in the target image; Performing image segmentation on the target image, and performing element region mapping processing on each segmented region to obtain the position coordinates corresponding to each element region in the target image, and establishing a mapping relationship between the element region in the target image and the row and column tuples; wherein the row and column tuples corresponding to each element region are consistent with the position index of each element in the target image; Based on the position index of the target element and the mapping relationship between the element region in the target image and the row and column tuples, the position coordinates corresponding to the element region indicated by the position index of the target element are determined to locate the target element.

[0005] In a second aspect, an embodiment of the present application provides an element positioning device, including: An acquisition module, used to acquire a target image where a target element to be located is located, and a position index of the target element, wherein the position index includes a row index and a column index, and the position index of each element in the target image is defined based on a position layout of each element in the target image; A segmentation module, used to perform image segmentation on the target image, and perform element region mapping processing on each segmented region to obtain the position coordinates corresponding to each element region in the target image, and establish a mapping relationship between the element region in the target image and the row and column tuples; wherein the row and column tuples corresponding to each element region are consistent with the position index of each element in the target image; A positioning module is used to determine the position coordinates corresponding to the element area indicated by the position index of the target element based on the position index of the target element and the mapping relationship between the element area in the target image and the row and column tuples, so as to locate the target element.

[0006] In a third aspect, an embodiment of the present application provides an electronic device, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the method described in the first aspect.

[0007] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, characterized in that a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.

[0008] In a fifth aspect, an embodiment of the present application provides a computer program product, which is stored in a computer-readable storage medium and is executed by at least one processor to implement the steps of the method described in the first aspect.

[0009] In an embodiment of the present application, firstly, a target image where the target element to be located is located, and a position index of the target element are obtained, the position index includes a row index and a column index, and the position index of each element in the target image is defined based on the position layout of each element in the target image, then the target image is segmented, and each segmented area is mapped to the element area to obtain the position coordinates corresponding to each element area in the target image, and a mapping relationship between the element area in the target image and the row and column tuple is established, wherein the row and column tuple corresponding to each element area is consistent with the position index of each element in the target image, and finally, based on the position index of the target element and the mapping relationship between the element area in the target image and the row and column tuple, the position coordinates corresponding to the element area indicated by the position index of the target element are determined to locate the target element. In an embodiment of the present application, the position coordinates of each element area in the image are obtained by segmenting the target image, and then the target element can be located by the position index input by the user, and the element is located based on the position coordinates of the element area. When the element rendering result remains unchanged, even if the attribute changes, there is no need to modify the code, and locating the target element by segmenting the image and the position index can be applied to different operating systems, and only one set of operating systems needs to be maintained, which is more convenient. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 is a flow chart of an element positioning method provided by an embodiment of the present application; Figure 2 is a schematic diagram of a target image provided by an embodiment of the present application; Figure 3 is a schematic diagram of indexes of various elements of a target image provided by an embodiment of the present application; Figure 4 is a schematic diagram of information of each region of a segmented target image provided by an embodiment of the present application; Figure 5 This is a schematic diagram of information of each region after removing the entire image region provided by an embodiment of the present application; Figure 6 is a schematic diagram of regional information of various elements provided by an embodiment of the present application; Figure 7 is a schematic diagram of the area information of each element provided by an embodiment of the present application after being divided into rows; Figure 8 is a schematic diagram of the sorted regional information rows of each element provided by an embodiment of the present application; Fig. 9 is a schematic diagram of the regional information of each element after being sorted in rows and columns provided by an embodiment of the present application; Fig.10 is a structural schematic diagram of an element positioning device provided by an embodiment of the present application; Fig.11 It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0011] The embodiments of the present application will be described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and should not be construed as limitations on the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.

[0012] The term "first" or "second" in the specification and claims of this application may include one or more of the features explicitly or implicitly. In the description of this application, unless otherwise specified, "plurality" means two or more. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally means that the objects connected before and after are in an "or" relationship.

[0013] like Figures 1 to 11 As shown, the embodiments of the present application provide an element positioning method, device and electronic device. The elements are positioned by image segmentation and position indexing to obtain the position information of the elements. The present application is applicable to situations where the boundaries between icon elements and adjacent elements are clear and the horizontal and vertical row and column layout rules of the elements in the figure are clear.

[0014] The meanings of the terms in the embodiments of this application are as follows: 1. Mobile terminal: Mobile terminal refers to computer equipment that can be used on the move, including mobile phones, tablet computers, notebooks, etc.

[0015] 2. User Interface: UI (User Interface) refers to the overall design of the software’s human-computer interaction, operation logic, and interface aesthetics. Here, it specifically refers to the user operation interface on mobile terminals.

[0016] 3. Pixels: Pixels are small squares in an image. These small squares have a clear position and assigned color values. The color and position of the small squares determine how the image appears. Pixels can be considered as indivisible units or elements in the entire image. Indivisible means that it cannot be cut into smaller units or elements. It exists as a small grid of a single color. Each dot matrix image contains a certain amount of pixels, which determine the size of the image on the screen.

[0017] 4. Element positioning: In UI automation testing, the user interface is composed of multiple element objects. The process of searching for elements through the name, label, value and other attributes of the element object is element positioning. Only after locating the element can subsequent automated testing operations such as input, click, and slide be performed.

[0018] 5. WebDriverAgent: WDA for short. It is a mobile testing framework launched by Facebook. It implements a WebDriver server on the iOS client. With this server, you can remotely control iOS devices for testing.

[0019] 6. SAM: A basic model in the field of visual segmentation, called Segment Anything Model, or SAM for short. This model mainly uses prompt engineering to train a large pre-trained model that performs segmentation based on prompts. The model has the potential to be applied in downstream segmentation tasks and can be combined with other visual tasks to form new solutions for other visual tasks. This work discusses the work from three perspectives: the selection of tasks for pre-training, model design, and data engine, and conducts a series of generalization experiments. This model has the strongest image segmentation Zero-shot capability at present, which is equivalent to GPT-3 in the field of computer vision.

[0020] 7. Android: Android is a free and open source operating system based on the Linux kernel (excluding GNU components), mainly used in mobile devices.

[0021] 8. SDK: Software Development Kit. A collection of development tools used by software development engineers to build application software for specific software packages, software frameworks, hardware platforms, operating systems, etc.

[0022] 9. UIAutomator: A testing framework released with the Android SDK. This testing framework provides a series of APIs that can be used to programmatically interact with Android apps, such as opening the settings menu, clicking, sliding, etc. It can also perform physical command operations on the device, such as rotating the device, obtaining the device resolution, etc.

[0023] 10. bbox: bbox is a method of expressing the coordinates of a rectangular area of ​​an image, that is, (x, y, w, h), where x represents the horizontal coordinate of the upper left corner, y represents the vertical coordinate of the upper left corner, w represents the width, and h represents the height.

[0024] like Figure 1 As shown, it is a flow chart of an element positioning method provided by an embodiment of the present application. Figure 1 As shown, the element positioning method may include the contents shown in steps 101 to 103.

[0025] In S101, a target image where a target element to be located is located and a position index of the target element are obtained. The position index includes a row index and a column index. The position index of each element in the target image is defined based on the position layout of each element in the target image.

[0026] The target image where the target element to be located is obtained can be obtained by automatically capturing the screen image. For example, the Android end can use the screenshot interface of the UIAutomator framework, and the iOS end can use the screenshot interface of the WebDriverAgent framework, such as Figure 2 As shown, the captured target image is a screenshot of the interface for logging into the first application.

[0027] In one example, Python language can be used for development, the version number can be Python3.8, the OpenCV library version used is opencv-python==4.8.0.76, and the imgread interface of the OpenCV library is used to load the screenshot to obtain the target image.

[0028] In this embodiment, the positioning Figure 2 The process of implementing element positioning is introduced by taking the button for switching username and password login as an example. The button for switching username and password login can be an email login button, that is Figure 2 The middle button at the bottom of the list. It can be determined based on the position index entered by the user.

[0029] Among them, the position index of the element can include two items, one is the row index, that is, the valid content of the image is divided into different rows, and the row numbers from top to bottom start from 0 and increase by 1 in sequence, and support reverse order, that is, the row number of the last row can also be represented by -1, the row number of the second to last row can also be represented by -2, and so on; the other is the column index, that is, each row of the image is divided into different columns, and the column numbers from left to right start from 0 and increase by 1 in sequence, and support reverse order, that is, the row number of the last row can also be represented by -1, the row number of the second to last row can also be represented by -2, and so on. In other words, the image can be divided into multiple rows and columns according to the position of each element in the image, and the position index of the element can be determined based on the row and column where the element is located. According to the above logic, the position index of the button for switching username and password login is: -1 row, 1 column, such as Figure 3 shown.

[0030] In S102, the target image is segmented, and each segmented area is mapped to an element area to obtain the position coordinates corresponding to each element area in the target image, and a mapping relationship between the element area in the target image and the row and column tuple is established.

[0031] The row and column tuples corresponding to each element region are consistent with the position index of each element in the target image.

[0032] In this embodiment, the target image is first segmented to obtain the areas and position coordinates of multiple segmented regions, and then the position coordinates of the element regions in each segmented region are determined, and a mapping relationship between the element regions and the row and column tuples is established based on the position coordinates of the element regions, that is, each element region corresponds to a row and a column. It should be noted that each element can be represented by a row and column tuple to represent its position in the target image.

[0033] In this embodiment, the SAM model may be used to segment the target image, or other methods may be used for segmentation, depending on the actual application, and this embodiment does not limit this.

[0034] In S103, based on the position index of the target element and the mapping relationship between the element region in the target image and the row and column tuples, the position coordinates corresponding to the element region indicated by the position index of the target element are determined to locate the target element.

[0035] That is to say, the position coordinates of the target element can be determined based on the position coordinates of each element area obtained by the above segmentation, the mapping relationship between the element area and the row and column tuples, and the position index of the target element input by the user, and then the target element can be located, so as to facilitate subsequent operations such as clicking and sliding on the target element.

[0036] In an embodiment of the present application, firstly, a target image where the target element to be located is located, and a position index of the target element are obtained, the position index includes a row index and a column index, and the position index of each element in the target image is defined based on the position layout of each element in the target image, then the target image is segmented, and each segmented area is mapped to the element area to obtain the position coordinates corresponding to each element area in the target image, and a mapping relationship between the element area in the target image and the row and column tuple is established, wherein the row and column tuple corresponding to each element area is consistent with the position index of each element in the target image, and finally, based on the position index of the target element and the mapping relationship between the element area in the target image and the row and column tuple, the position coordinates corresponding to the element area indicated by the position index of the target element are determined to locate the target element. In an embodiment of the present application, the position coordinates of each element area in the image are obtained by segmenting the target image, and then the target element can be located by the position index input by the user, and the element is located based on the position coordinates of the element area. When the element rendering result remains unchanged, even if the attribute changes, there is no need to modify the code, and locating the target element by segmenting the image and the position index can be applied to different operating systems, and only one set of operating systems needs to be maintained, which is more convenient.

[0037] In a possible implementation of the present application, a target image is segmented, and each segmented area is mapped to an element area to obtain the position coordinates corresponding to each element area in the target image, and a mapping relationship between the element area in the target image and the row and column tuples is established. This may include: inputting the target image into a SAM model for image segmentation to obtain a segmentation result, wherein the segmentation result includes the area and position coordinates of each segmented area; performing element area screening on each segmented area in the segmentation result to obtain the area and position coordinates corresponding to each element area; classifying and sorting the screened element areas according to the rows and columns defined in the target image, and establishing a mapping relationship between the element area of ​​the target image and the row and column tuples.

[0038] In this embodiment, the target image is input into the SAM model for image segmentation, and information such as the area and bbox coordinates of multiple segmented regions can be obtained. Then, through screening, classification, sorting and other processing, a mapping relationship between each element area in the target image and the row and column binary tuples is established.

[0039] This embodiment uses the SAM model to segment the target image as an example. First, the sam_model_registry function in the SAM code is used to load the MetaAI official pre-trained model, using the sam_vit_b_01ec64 model. Then, a segmenter is initialized using the SamAutomaticMaskGenerator function in the SAM code. Finally, the target image is input into the segmenter to generate the segmentation result, that is, the area and position coordinates of each segmented area. Then, the element area of ​​each segmented area in the segmentation result is screened, and the area and position coordinates corresponding to the non-element area are deleted. The area and position coordinates corresponding to each element area can be obtained. After classification and sorting, the mapping relationship between each element area of ​​the target image and the row and column binary can be obtained. SAM adopts an encoder-decoder structure, in which the encoder part is composed of multiple convolutional layers and pooling layers for extracting image features. Each convolutional layer usually includes operations such as convolution kernels, activation functions, and batch normalization, which are used to extract features and reduce the dimension of the input image. The pooling layer is used to downsample the feature map to reduce the amount of calculation and memory consumption; the decoder part consists of multiple deconvolution layers and upsampling layers. Each deconvolution layer usually includes operations such as deconvolution kernel, activation function and batch normalization, which are used to upsample and fuse the feature map. The upsampling layer is used to upsample the feature map to restore the original image size and generate the segmentation result. The output segmentation result is a list. Each list object contains information such as the area of ​​the segmented area, bbox coordinates, etc. Figure 4 shown.

[0040] In a possible implementation of the present application, each segmented area in the segmentation result is screened for elemental areas to obtain the area and position coordinates corresponding to each elemental area, which may include: deleting the entire image area in the segmentation result based on the area of ​​each segmented area; the entire image area refers to the area with the largest area in the segmentation result; based on the area and position coordinates of each segmented area after deleting the entire image area, screening out each elemental area in the target image, and obtaining the area and position coordinates corresponding to each elemental area.

[0041] Since each segmentation result must include the area of ​​the complete image, we can first remove the area of ​​the entire image. Figure 4The area with the largest area is the area of ​​the entire image, that is, the whole image area. This item has no effect on the subsequent positioning of elements in the image. In order to reduce the amount of calculation, this item can be removed first, and then the area and position coordinates of each segmented area after deleting the whole image area can be used to filter out the various element areas in the target image, as well as the area and position coordinates corresponding to each element area. By filtering out the element areas and their position coordinates, the positions of the various elements in the overall image can be clearly reflected, so as to correspond to the row and column tuples, and then quickly locate the target elements. In this embodiment, Python code can be used for comparison Figure 4 The size of the area field in the comparison, that is, the size of the area, is 2748062, which is the largest area. After removing it, the remaining segmented areas are as follows Figure 5 shown.

[0042] In a possible implementation of the present application, based on the area and position coordinates of each segmented area after deleting the entire image area, each element area in the target image is screened out, and the area and position coordinates corresponding to each element area are obtained, which may include: comparing the position coordinates of each segmented area in pairs to determine whether there is an overlapping relationship between each two segmented areas; deleting the segmented area with a smaller area among the two segmented areas with an overlapping relationship, using all retained segmented areas as the screened out element areas, and obtaining the area and position coordinates of each element area; an element area refers to the area where any element in the target image is located.

[0043] After deleting the entire image area in the above embodiment, each element area and the corresponding area and position coordinates can be further screened out. Figure 5 The bbox coordinates of all segmented areas in the image are compared pairwise to determine the overlapping relationship between different areas. Then, the segmented area with the smaller area in the two overlapping segmented areas is deleted, and all the retained segmented areas are used as the screened element areas, and the area and position coordinates of each element area are obtained.

[0044] The comparison process of the bbox coordinates of any two segmented areas is as follows: Figure 2 Take the icon in the lower right corner as an example, the icon area corresponds to Figure 5The area in the figure is 5645, and the bbox coordinates are [671, 1995, 88, 88]. The five-pointed star area in the icon corresponds to the area in the figure above with an area of ​​1627, and the bbox coordinates are [692, 2015, 41, 50], the four numbers of the bbox coordinates are the horizontal coordinate x, the vertical coordinate y, the width w, and the height h. The horizontal coordinate 692 of the upper left corner and the horizontal coordinate 692+41=733 of the lower right corner of the five-pointed star area are both located between the horizontal coordinate 671 of the upper left corner and the horizontal coordinate 671+88=759 of the lower right corner of the circular icon. The vertical coordinate 2015 of the upper left corner and the vertical coordinate 2015+50=2065 of the lower right corner of the five-pointed star area are both located between the vertical coordinate 1995 of the upper left corner and the vertical coordinate 1995+88=2083 of the lower right corner of the circular icon area. Therefore, it is judged that the area with an area of ​​1627 is included in the area with an area of ​​5645. The area with an area of ​​1627 is removed and the area with an area of ​​5645 is retained.

[0045] By comparing the overlapping relationship between any two regions one by one in this embodiment, 20 regions can be obtained, such as Figure 6 As shown, these are the areas corresponding to each element in the target image.

[0046] This embodiment can delete invalid region coordinates by determining the region corresponding to each element in the target image, thereby facilitating subsequent positioning of each element and reducing the amount of calculation.

[0047] In a possible implementation of the present application, each selected element region is classified and sorted according to the rows and columns defined in the target image, and a mapping relationship between the element region of the target image and the row and column tuple is established, which may include: based on the vertical coordinate in the position coordinates of each element region in the target image, each element region in the target image is classified according to the row, and each classified row is sorted; based on the horizontal coordinate in the position coordinates of each element region in each row, each element region in each row is sorted according to the column; based on the row and column where each element region in the target image is located after the classification and sorting, a mapping relationship between the element region of the target image and the row and column tuple is established.

[0048] After the position coordinates of each element region are determined in the above embodiment, each region can be classified and sorted in rows and columns to obtain the row and column where each element region in the target image is located, that is, the relative position relationship of each element region in the target image can be determined, so as to establish a mapping relationship between the element region and the row and column tuple of the target image. The position relationship between elements is determined by the row and column where the element region is located in the target image, so that the positioning of the element has nothing to do with its attributes. Even if the attributes of the element change, its relative position relationship will not change, making the positioning more accurate. When the attributes of the element change, there is no need to update the code, thereby reducing maintenance costs.

[0049] In a possible implementation of the present application, based on the vertical coordinate in the position coordinate of each element area in the target image, each element area in the target image is classified by row, which may include: traversing all element areas in the target image, and determining the first selected element area as an area in the first row; for the next element area selected in sequence, comparing the current element area with each area in each row, if the current element area has an overlapping relationship with any area in the current row in terms of the vertical coordinate, determining that the current element area belongs to the current row, if the current element area does not have an overlapping relationship with all areas in all rows in terms of the vertical coordinate, determining that the current element area belongs to a new row.

[0050] In this embodiment, an element region can be selected first and determined as the region of the first row, and then the element region is compared with all other element regions. If there is an overlapping relationship between the vertical coordinates, they belong to the same row. If there is no overlapping relationship between the vertical coordinates, they do not belong to the row, and the element region belongs to a new row. The vertical coordinates are compared in sequence to determine multiple rows of element regions. Then, the rows can be sorted based on the size of the vertical coordinates of each row. Based on the same principle, the element region of each row can be divided into columns based on the horizontal coordinates. By sorting the rows and columns, the positional relationship of each element in the target image can be determined.

[0051] In one example, you can use Python code to Figure 6 The 20 segmented areas in the table are classified by row. The specific implementation logic is as follows: First, Figure 6 The first region of the segmented region is taken as a region of the first row, and then Figure 6 The second region of the obtained segmented region is denoted by region X. Region X is compared with each region in the current row. If region X has an overlapping relationship with any region in the current row in terms of ordinate, then region X belongs to the current row and is put into the list of the current row. If region X has no overlapping relationship with all regions in all rows in terms of ordinate, then region X is put into the list of a new row and then Figure 6 The third area of ​​the segmented area is referred to as the Y area, and the same comparison logic is performed on the X area until Figure 6 The process ends after the last area of ​​the obtained segmented area completes the row classification. Use Python code to compare the vertical coordinates to determine whether there is an overlapping relationship between the vertical coordinates of the two segmented areas. For example, the top and bottom vertical coordinates of one segmented area are top1 and bottom1 respectively, and the top and bottom vertical coordinates of the other segmented area are top2 and bottom2 respectively. The logical expression for determining whether the vertical coordinates of the two areas overlap is: top2<=top1<=bottom2 or top2<=bottom1<=bottom2 or top1<=top2<=bottom1 or top1<=bottom2<=bottom1. If the return result of this expression is True, the vertical coordinates of the two areas overlap, otherwise not. After the above logic is compared, Figure 6 The results of classifying the 20 regions by row are as follows Figure 7 As shown, Figure 7 The 20 segmentation areas are divided into 7 rows.

[0052] Then sort the classified rows, you can use Python code to Figure 7 The top vertical coordinates of the first area of ​​each row are in ascending order. Figure 7 All rows are sorted, and after sorting, all rows are in top-down order, such as Figure 8 shown.

[0053] Then sort the columns of each row, and you can use Python code to Figure 8 Each row of the output is sorted in ascending order according to the horizontal coordinate on the left side of each area, that is, the column sorting of each row is completed. After sorting, the areas of each row are arranged from left to right according to their positions in the target image, such as Fig. 9 shown.

[0054] The coordinates of the target element are obtained by the position index entered by the user, that is, the result of element positioning. For example, the position index of the button for switching username and password login entered by the user is: -1 row, 1 column. Fig. 9After the output columns are sorted, this system uses Python code to obtain the result that the -1th row is [{'area': 6148, 'bbox': [373, 1995, 88,88]}, {'area': 6197, 'bbox': [522, 1995, 87, 88]}, {'area': 5645, 'bbox':[671, 1995, 88, 88]}], where the 1st column is {'area': 6197, 'bbox': [522, 1995, 87, 88]}, that is, the bbox coordinates of the target element's positioning result are [522, 1995, 87, 88].

[0055] The present application also provides an element positioning device. Fig.10 As shown in FIG. 1 , it is a schematic diagram of the structure of an element positioning device provided in an embodiment of the present application. Fig.10 As shown, the element positioning device may include: an acquisition module 1001 , a segmentation module 1002 and a positioning module 1003 .

[0056] Among them, the acquisition module 1001 is used to obtain the target image where the target element to be located is located, and the position index of the target element, the position index includes a row index and a column index, and the position index of each element in the target image is defined based on the position layout of each element in the target image; the segmentation module 1002 is used to perform image segmentation on the target image, and map each segmented area to the element area to obtain the position coordinates corresponding to each element area in the target image, and establish a mapping relationship between the element area in the target image and the row and column tuples; wherein the row and column tuples corresponding to each element area are consistent with the position index of each element in the target image; the positioning module 1003 is used to determine the position coordinates corresponding to the element area indicated by the position index of the target element based on the position index of the target element and the mapping relationship between the element area in the target image and the row and column tuples, so as to locate the target element.

[0057] In an embodiment of the present application, first, the acquisition module 1001 acquires the target image where the target element to be located is located, and the position index of the target element, the position index includes a row index and a column index, and the position index of each element in the target image is defined based on the position layout of each element in the target image, and then the segmentation module 1002 performs image segmentation on the target image, and performs element region mapping processing on each segmented area to obtain the position coordinates corresponding to each element area in the target image, and establishes a mapping relationship between the element area in the target image and the row and column binary tuple, wherein the row and column binary tuple corresponding to each element area is consistent with the position index of each element in the target image, and finally the positioning module 1003 determines the position coordinates corresponding to the element area indicated by the position index of the target element based on the position index of the target element and the mapping relationship between the element area in the target image and the row and column binary tuple to locate the target element. The embodiment of the present application obtains the position coordinates of each element area in the image by performing image segmentation on the target image, and then locates the target element through the position index input by the user. The element is located based on the position coordinates of the element area. When the element rendering result remains unchanged, there is no need to modify the code even if the attributes change. In addition, locating the target element by segmenting the image and the position index can be applicable to different operating systems, and only one set of operating system needs to be maintained, which is more convenient.

[0058] In a possible implementation of the present application, the segmentation module 1002 is used to: input the target image into the SAM model for image segmentation to obtain a segmentation result, wherein the segmentation result includes the area and position coordinates of each segmented region; perform element region screening on each segmented region in the segmentation result to obtain the area and position coordinates corresponding to each element region; classify and sort the screened element regions according to the rows and columns defined in the target image, and establish a mapping relationship between the element region of the target image and the row and column tuples.

[0059] In a possible implementation of the present application, the segmentation module 1002 is used to: delete the whole image area in the segmentation result based on the area of ​​each segmented area; the whole image area refers to the area with the largest area in the segmentation result; based on the area and position coordinates of each segmented area after deleting the whole image area, filter out each element area in the target image, and obtain the area and position coordinates corresponding to each element area.

[0060] In a possible implementation of the present application, the segmentation module 1002 is used to: compare the position coordinates of each segmentation area in pairs to determine whether there is an overlapping relationship between each two segmentation areas; delete the segmentation area with a smaller area among the two segmentation areas with an overlapping relationship, and use all retained segmentation areas as the screened element areas, and obtain the area and position coordinates of each element area; the element area refers to the area where any element in the target image is located.

[0061] In a possible implementation of the present application, the positioning module 1003 is used to: classify each element region in the target image according to rows based on the vertical coordinates in the position coordinates of each element region in the target image, and sort the classified rows; sort each element region in each row according to columns based on the horizontal coordinates in the position coordinates of each element region in each row; and establish a mapping relationship between the element region of the target image and the row and column tuple based on the row and column where each element region in the target image is located after classification and sorting.

[0062] In a possible implementation of the present application, the positioning module 1003 is used to: traverse all element regions in the target image, and determine the first selected element region as a region in the first row; for the next element region selected in sequence, compare the current element region with each region in each row, and if the current element region has an overlapping relationship in the vertical coordinate with any region in the current row, determine that the current element region belongs to the current row; if the current element region does not have an overlapping relationship in the vertical coordinate with all regions in all rows, determine that the current element region belongs to a new row.

[0063] The element positioning device in the embodiment of the present application can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. The non-mobile electronic device can be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), a teller machine or a self-service machine, etc., which is not specifically limited in the embodiment of the present application.

[0064] The element positioning device in the embodiment of the present application may be a device having an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.

[0065] The element positioning device provided in the embodiment of the present application can achieve Figures 1 to 9 The various processes implemented in the method embodiment achieve the same technical effect and will not be described again here to avoid repetition.

[0066] Fig.11 A schematic diagram of the hardware structure of an electronic device for implementing each embodiment of the present application is provided below. The electronic device 1100 includes but is not limited to: a radio frequency unit 1101, a network module 1102, an audio output unit 1103, an input unit 1104, a sensor 1105, a display unit 1106, a user input unit 1107, an interface unit 1108, a memory 1109, a processor 1110, and a power supply 1111. Those skilled in the art will appreciate that Fig.11 The electronic device structure shown in the figure does not constitute a limitation on the electronic device, and the electronic device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently. In the embodiments of the present application, the electronic device includes but is not limited to a mobile phone, a tablet computer, a laptop computer, a PDA, a vehicle-mounted terminal, a wearable device, and a pedometer.

[0067] Among them, the processor 1110 can be used to: obtain a target image where the target element to be located is located, and a position index of the target element, the position index includes a row index and a column index, and the position index of each element in the target image is defined based on the position layout of each element in the target image; perform image segmentation on the target image, and perform element region mapping processing on each segmented area to obtain the position coordinates corresponding to each element area in the target image, and establish a mapping relationship between the element area in the target image and the row and column tuples; wherein the row and column tuples corresponding to each element area are consistent with the position index of each element in the target image; based on the position index of the target element, and the mapping relationship between the element area in the target image and the row and column tuples, determine the position coordinates corresponding to the element area indicated by the position index of the target element to locate the target element.

[0068] In an embodiment of the present application, firstly, a target image where the target element to be located is located, and a position index of the target element are obtained, the position index includes a row index and a column index, and the position index of each element in the target image is defined based on the position layout of each element in the target image, then the target image is segmented, and each segmented area is mapped to the element area to obtain the position coordinates corresponding to each element area in the target image, and a mapping relationship between the element area in the target image and the row and column tuple is established, wherein the row and column tuple corresponding to each element area is consistent with the position index of each element in the target image, and finally, based on the position index of the target element and the mapping relationship between the element area in the target image and the row and column tuple, the position coordinates corresponding to the element area indicated by the position index of the target element are determined to locate the target element. In an embodiment of the present application, the position coordinates of each element area in the image are obtained by segmenting the target image, and then the target element can be located by the position index input by the user, and the element is located based on the position coordinates of the element area. When the element rendering result remains unchanged, even if the attribute changes, there is no need to modify the code, and locating the target element by segmenting the image and the position index can be applied to different operating systems, and only one set of operating systems needs to be maintained, which is more convenient.

[0069] It should be understood that in the embodiment of the present application, the radio frequency unit 1101 can be used for receiving and sending signals during information transmission or calls. Specifically, after receiving downlink data from the base station, it is sent to the processor 1110 for processing; in addition, the uplink data is sent to the base station. Generally, the radio frequency unit 1101 includes but is not limited to an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier, a duplexer, etc. In addition, the radio frequency unit 1101 can also communicate with the network and other devices through a wireless communication system.

[0070] The electronic device provides users with wireless broadband Internet access through the network module 1102, such as helping users to send and receive emails, browse web pages, and access streaming media.

[0071] The audio output unit 1103 can convert the audio data received by the RF unit 1101 or the network module 1102 or stored in the memory 1109 into an audio signal and output it as sound. Moreover, the audio output unit 1103 can also provide audio output related to a specific function performed by the electronic device 1100 (for example, a call signal reception sound, a message reception sound, etc.). The audio output unit 1103 includes a speaker, a buzzer, a receiver, etc.

[0072] The input unit 1104 is used to receive audio or video signals. The input unit 1104 may include a graphics processor (GPU) 11041 and a microphone 11042. The graphics processor 11041 processes the image data of a static picture or video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The processed image frame can be displayed on the display unit 1106. The image frame processed by the graphics processor 11041 can be stored in the memory 1109 (or other storage medium) or sent via the radio frequency unit 1101 or the network module 1102. The microphone 11042 can receive sound and can process such sound into audio data. The processed audio data can be converted into a format output that can be sent to a mobile communication base station via the radio frequency unit 1101 in the case of a telephone call mode.

[0073] The electronic device 1100 also includes at least one sensor 1105, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor includes an ambient light sensor and a proximity sensor, wherein the ambient light sensor can adjust the brightness of the display panel 11061 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 11061 and / or the backlight when the electronic device 1100 is moved to the ear. As a type of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in each direction (generally three axes), and can detect the magnitude and direction of gravity when stationary, which can be used to identify the posture of the electronic device (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; the sensor 1105 can also include a fingerprint sensor, a pressure sensor, an iris sensor, a molecular sensor, a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, etc., which will not be repeated here.

[0074] The display unit 1106 is used to display information input by the user or information provided to the user. The display unit 1106 may include a display panel 11061, which may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like.

[0075] The user input unit 1107 can be used to receive input digital or character information, and generate key signal input related to user settings and function control of the electronic device. Specifically, the user input unit 1107 includes a touch panel 11071 and other input devices 11072. The touch panel 11071, also known as a touch screen, can collect user touch operations on or near it (such as operations performed by users using fingers, styluses, or any other suitable objects or accessories on or near the touch panel 11071). The touch panel 11071 may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the user's touch orientation, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact point coordinates, and then sends it to the processor 1110, receives the command sent by the processor 1110 and executes it. In addition, the touch panel 11071 can be implemented in various types such as resistive, capacitive, infrared, and surface acoustic waves. In addition to the touch panel 11071, the user input unit 1107 may also include other input devices 11072. Specifically, other input devices 11072 may include but are not limited to physical keyboards, function keys (such as volume control keys, switch keys, etc.), trackballs, mice, and joysticks, which will not be repeated here.

[0076] Furthermore, the touch panel 11071 may be overlaid on the display panel 11061. When the touch panel 11071 detects a touch operation on or near it, it is transmitted to the processor 1110 to determine the type of touch event. Then, the processor 1110 provides corresponding visual output on the display panel 11061 according to the type of touch event. Fig.11 In the figure, the touch panel 11071 and the display panel 11061 are two independent components to realize the input and output functions of the electronic device. However, in some embodiments, the touch panel 11071 and the display panel 11061 can be integrated to realize the input and output functions of the electronic device, which is not limited here.

[0077] The interface unit 1108 is an interface for connecting an external device to the electronic device 1100. For example, the external device may include a wired or wireless headset port, an external power supply (or battery charger) port, a wired or wireless data port, a memory card port, a port for connecting a device with an identification module, an audio input / output (I / O) port, a video I / O port, a headphone port, etc. The interface unit 1108 may be used to receive input (e.g., data information, power, etc.) from an external device and transmit the received input to one or more elements within the electronic device 1100 or may be used to transmit data between the electronic device 1100 and an external device.

[0078] The memory 1109 can be used to store software programs and various data. The memory 1109 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area can store data created according to the use of the mobile phone (such as audio data, a phone book, etc.), etc. In addition, the memory 1109 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0079] The processor 1110 is the control center of the electronic device. It uses various interfaces and lines to connect various parts of the entire electronic device. By running or executing software programs and / or modules stored in the memory 1109 and calling data stored in the memory 1109, it performs various functions of the electronic device and processes data, thereby monitoring the electronic device as a whole. The processor 1110 may include one or more processing units; preferably, the processor 1110 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 1110.

[0080] The electronic device 1100 may also include a power supply 1111 (such as a battery) for supplying power to various components. Preferably, the power supply 1111 may be logically connected to the processor 1110 via a power management system, thereby implementing functions such as charging, discharging, and power consumption management through the power management system.

[0081] In addition, the electronic device 1100 includes some functional modules not shown, which will not be described in detail here.

[0082] An embodiment of the present application also provides an electronic device, including a processor 1110, a memory 1109, and a computer program stored in the memory 1109 and executable on the processor 1110. When the computer program is executed by the processor 1110, each process of the above-mentioned element positioning method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.

[0083] The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, each process of the above-mentioned element positioning method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it is not repeated here. The computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0084] An embodiment of the present application provides a computer program product, which is stored in a computer-readable storage medium. The computer program product is executed by at least one processor to implement the various processes of the above-mentioned element positioning method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0085] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned element positioning method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0086] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.

[0087] It should be noted that, in this article, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0088] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0089] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.

Claims

1. A method for locating an element, characterized in that: include: Acquire a target image where the target element to be located is located, and a position index of the target element, wherein the position index includes a row index and a column index, and the position index of each element in the target image is defined based on the position layout of each element in the target image; Performing image segmentation on the target image, and performing element region mapping processing on each segmented region to obtain the position coordinates corresponding to each element region in the target image, and establishing a mapping relationship between the element region in the target image and the row and column tuples; wherein the row and column tuples corresponding to each element region are consistent with the position index of each element in the target image; Based on the position index of the target element and the mapping relationship between the element region in the target image and the row and column tuples, the position coordinates corresponding to the element region indicated by the position index of the target element are determined to locate the target element.

2. The method according to claim 1, characterized in that The step of segmenting the target image and performing element region mapping processing on each segmented region to obtain position coordinates corresponding to each element region in the target image and establishing a mapping relationship between the element region in the target image and the row and column tuples includes: Input the target image into the SAM model to perform image segmentation to obtain a segmentation result, wherein the segmentation result includes the area and position coordinates of each segmented area; Screening the element regions of each segmented region in the segmentation result to obtain the area and position coordinates corresponding to each element region; The selected element regions are classified and sorted according to the rows and columns defined in the target image, and a mapping relationship between the element regions of the target image and the row and column tuples is established.

3. The method according to claim 2, characterized in that The screening of element regions for each segmented region in the segmentation result to obtain the area and position coordinates corresponding to each element region includes: Based on the areas of the respective segmented areas, the entire image area in the segmentation result is deleted; the entire image area refers to the area with the largest area in the segmentation result; Based on the area and position coordinates of each segmented area after deleting the entire image area, each element area in the target image is screened out, and the area and position coordinates corresponding to each element area are obtained.

4. The method according to claim 3, characterized in that The method of screening out each element region in the target image based on the area and position coordinates of each segmented region after deleting the entire image region, and obtaining the area and position coordinates corresponding to each element region, includes: Compare the position coordinates of each segmented area in pairs to determine whether there is an overlapping relationship between any two segmented areas; The segmented area with the smaller area among the two overlapping segmented areas is deleted, and all the retained segmented areas are used as the screened element areas, and the area and position coordinates of each element area are obtained; the element area refers to the area where any element in the target image is located.

5. The method according to claim 2, characterized in that: Classifying and sorting the selected element regions according to the rows and columns defined in the target image, and establishing a mapping relationship between the element regions of the target image and the row and column tuples, including: Based on the ordinate in the position coordinate of each element region in the target image, classifying each element region in the target image according to rows, and sorting each row obtained by classification; Based on the horizontal coordinates in the position coordinates of each element region of each row, each element region of each row is sorted by column; Based on the row and column where each element region in the target image is located after the classification and sorting processing, a mapping relationship between the element region of the target image and the row and column tuple is established.

6. The method according to claim 5, characterized in that The classifying the element regions in the target image according to rows based on the ordinates in the position coordinates of the element regions in the target image includes: Traverse all element regions in the target image, and determine the first selected element region as a region in the first row; for the next element region selected in sequence, compare the current element region with each region in each row, and if the current element region has an overlapping relationship in the vertical coordinate with any region in the current row, determine that the current element region belongs to the current row; if the current element region does not have an overlapping relationship in the vertical coordinate with all regions in all rows, determine that the current element region belongs to a new row.

7. An element positioning device, characterized in that: include: An acquisition module, used to acquire a target image where a target element to be located is located, and a position index of the target element, wherein the position index includes a row index and a column index, and the position index of each element in the target image is defined based on a position layout of each element in the target image; A segmentation module, used to perform image segmentation on the target image, and perform element region mapping processing on each segmented region to obtain the position coordinates corresponding to each element region in the target image, and establish a mapping relationship between the element region in the target image and the row and column tuples; wherein the row and column tuples corresponding to each element region are consistent with the position index of each element in the target image; A positioning module is used to determine the position coordinates corresponding to the element area indicated by the position index of the target element based on the position index of the target element and the mapping relationship between the element area in the target image and the row and column tuples, so as to locate the target element.

8. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the steps of the method according to any one of claims 1 to 6 when executed by the processor.

9. A computer-readable storage medium, characterized in that: include: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product, characterized in that The computer program product is stored in a computer-readable storage medium, and the computer program product is executed by at least one processor to implement the steps of the method according to any one of claims 1 to 6.

Citation Information

Cited By

  • Vector element arrangement method and device, electronic equipment and storage medium

    CN122086892A

  • Arrangement method and device of vector elements, electronic equipment and storage medium

    CN122086892B