System and method for automatically testing confused APP based on visual identification
Through visual recognition technology, combined with manifold learning and Kalman filtering, the problems of low positioning accuracy of traditional controls and poor adaptability of dynamic interfaces are solved, efficient and stable control recognition and operation execution are achieved, and the accuracy and reliability of automated testing are improved.
Patent Information
- Application Number
- CN202510325672.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-04
AI Technical Summary
Traditional control positioning methods are difficult to ensure accuracy in complex or dynamically changing user interfaces, dynamic elements affect test stability, low image preprocessing efficiency, and inaccurate execution of automated operations.
An obfuscated APP automated testing system based on visual recognition is adopted, including image input, preprocessing, manifold learning, information geometric matching, dynamic element processing and automated operation execution modules. Dynamic changes are processed through Kalman filtering and partial differential equations, combining manifold learning and information geometry optimization matching accuracy.
It improves control positioning accuracy and dynamic element adaptability, improves image preprocessing efficiency and stability and reliability of automated operations, and ensures efficient testing under complex interfaces.
Smart Images

Figure CN120256299A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of visual recognition, and specifically to a system and method for automated testing of obfuscated APPs based on visual recognition. Background Art
[0002] First of all, traditional control positioning methods usually rely on image matching technology, which is effective in some simple scenarios, but for complex or dynamically changing user interfaces, the accuracy is often difficult to guarantee. Existing technologies often determine the control position by directly comparing the target image with the screen screenshot. This method ignores the geometric deformation of the image and the dynamic state changes of the control. As a result, when the control position shifts, or the interface displays different contents in different states, the matching accuracy and stability cannot be guaranteed.
[0003] Secondly, most existing technologies ignore the impact of dynamic elements on the testing process. Many existing automated testing systems are only effective in static scenarios and cannot cope with the dynamic changes in the position and state of controls in the interface. For example, when dynamic controls such as scroll bars and pop-up windows appear, traditional methods cannot adapt to these changes in real time, resulting in control positioning errors and further affecting the execution of subsequent automated operations. This also makes the existing testing systems often inefficient and unstable when dealing with dynamic applications.
[0004] Furthermore, traditional image preprocessing methods are single and have low efficiency. Most existing image preprocessing is limited to basic denoising operations and often ignores optimization steps such as image cropping and contrast enhancement. The lack of effective image quality improvement methods leads to a decrease in accuracy during subsequent image matching. Moreover, due to the lack of integration of multiple processing technologies, existing methods may still have problems such as image blurring and noise interference in complex images, affecting the accuracy and stability of automated testing.
[0005] Finally, traditional automated operation execution systems have many assumptions and fixed patterns. Many systems only simulate basic operations such as clicking and input, but cannot flexibly adapt to the dynamic changes of controls and even lack the ability to respond immediately to control state changes; Therefore, the present invention proposes a system and method for automated testing of obfuscated APPs based on visual recognition to solve the deficiencies of the existing technologies. Summary of the Invention
[0006] Aiming at the deficiencies of the existing technologies, the present invention provides a system and method for automated testing of obfuscated APPs based on visual recognition, which solves the problems of low accuracy of traditional image matching, poor adaptability to dynamic controls, low efficiency of image preprocessing, and inaccurate execution of automated operations.
[0007] To achieve the above object, the present invention is implemented by the following technical solutions: An automated testing system for obfuscated APPs based on visual recognition, characterized by comprising: An image input module for receiving a target control image and an APP screen image; An image preprocessing module for denoising, cropping, and contrast enhancement processing on the received images; A manifold learning module for performing dimensionality reduction on the preprocessed images, extracting the geometric features of the images, and performing feature embedding on the control image and the background image in the low-dimensional manifold space; An information geometry matching module for calculating the similarity between the control image and the background image according to the geometric features of the images, and optimizing the image matching accuracy based on mutual information and KL divergence; An image transformation module for performing diffusion processing on the images based on partial differential equations and performing affine transformation according to the changes of the images to ensure the geometric stability of the images; A dynamic element processing module for dynamically predicting and correcting the position of the control during the image matching process through Kalman filtering; An automated operation execution module for performing corresponding automated operations after successful control positioning; A report generation module for recording and generating the operation results and test reports of the test process.
[0008] The present invention also provides an automated testing method for obfuscated APPs based on visual recognition, comprising the following steps: S1. Receive a target control image and a screen image to be recognized; S2. Preprocess the input images, including denoising, cropping, and contrast enhancement operations; S3. Use manifold learning to perform dimensionality reduction on the preprocessed images and extract the low-dimensional geometric features of the images; S4. Calculate the similarity between the control image and the screen image based on information geometry, and optimize the matching accuracy by calculating the mutual information value and KL divergence; S5. Perform diffusion processing on the images through partial differential equations and apply affine transformation to correct the geometric changes of the images; S6. Dynamically predict and correct the error of the control position through Kalman filtering; S7. Perform automated test operations after successful control positioning; S8. Record the test results and generate test reports The present invention provides an automated testing system and method for obfuscated APPs based on visual recognition. It has the following beneficial effects: 1. The present invention adopts a technical solution that combines manifold learning and information geometry, achieving higher control positioning accuracy. Compared with the method that only relies on image matching in the prior art, the present invention can effectively solve the problem of inaccurate positioning of complex interfaces and dynamic controls in traditional methods by extracting the low-dimensional geometric features of images and calculating the information geometric similarity. Through these innovative means, the positioning accuracy is greatly improved, especially suitable for complex and variable UI interfaces.
[0009] 2. The present invention adopts a technical solution that uses Kalman filtering and partial differential equations to process dynamic elements, achieving stronger adaptability to dynamic elements. Compared with the limitations of static image processing in the prior art, the present invention can predict the position change of controls in real time and correct errors, effectively solving the problem that traditional methods cannot respond to the dynamic changes of controls in real time. This enables the system to still maintain efficient and stable control recognition and operation execution capabilities when facing dynamic interfaces.
[0010] 3. The present invention adopts multiple image preprocessing operations such as denoising, cropping, and contrast enhancement, achieving higher processing efficiency and image quality. Compared with the single image processing means in the prior art, the present invention solves the problems of slow image processing speed and poor effect in the existing solutions by integrating multiple technologies during the preprocessing process. The preprocessed image is clearer, ensuring the accuracy of subsequent image matching and control positioning.
[0011] 4. The present invention realizes precise interactive operations on controls through precise control positioning and automated operation execution technical solutions. Compared with the simple operation simulation method in the prior art, the present invention can avoid the problems of operation deviation and error in traditional automated testing through real-time updating of control positions and dynamic prediction technologies. The precise execution of automated operations greatly improves the stability and reliability of testing and reduces the need for human intervention. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 is the system architecture diagram of the present invention; Figure 2 is the architecture diagram of the manifold learning module of the present invention; Figure 3 is the architecture diagram of the information geometry matching module of the present invention; Figure 4 is the architecture diagram of the dynamic element processing module of the present invention; Figure 5 is the method flow chart of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0013] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0014] Please refer to Figures 1-4 , the embodiment of the present invention provides a post-confusion APP automated testing system based on visual recognition, including: an image input module for receiving a target control image and an APP screen image; In this embodiment, the image input module receives the control image and the screen image in various ways. The user can manually capture the control on the current interface through the system interface or select the image of the target control from a preset image library. The control image can be a screenshot of interface elements such as buttons and icons, or a screenshot of the entire control area; the screen image is a complete screenshot of the current APP interface. The image input module can efficiently capture and transfer these image data to ensure that the subsequent processing module can receive standardized input.
[0015] In some embodiments, the image input module provides the system with a screen capture of the current APP for subsequent image preprocessing. To ensure the adaptability and efficiency of the system, the image input module can automatically identify and extract the control elements in the screen, and at the same time support manual selection of specific control images for input. This flexibility improves the applicability of the system when processing different APP interfaces, especially when the interface changes frequently.
[0016] In specific implementation, the image input module is closely connected to subsequent image preprocessing modules, manifold learning modules, etc. After receiving the control image and the screen image, the image input module transfers them to the image preprocessing module, which performs processing such as denoising, cropping, and contrast enhancement on the images to ensure that the image quality meets the requirements. The processed images will be further transferred to the manifold learning module for dimensionality reduction and feature extraction. The output of this module is a low-dimensional feature vector, providing the necessary geometric features for subsequent image matching.
[0017] Specifically, the image input module receives the control image and the screen image input by the user through an interface. In some embodiments, this module may be equipped with an image capture function that can automatically detect the control elements that appear on the screen and use the images of these controls and the entire screen image as input data. The user can also choose to manually upload the control image, and the system performs subsequent test operations based on these images.
[0018] As an option, the image input module can automatically identify the control areas in the image and extract them. This process can be completed through a control detection algorithm based on image recognition technology. In one possible implementation, the system identifies the control areas in the image by the preset shapes of common UI elements such as sliders and buttons. In this way, the image input module can provide more accurate input for subsequent image matching and control positioning.
[0019] In some embodiments, the image input module performs preliminary cleaning on the input image to avoid the damaged or blurred images affecting the test results during the subsequent processing of the system. Specifically, the module performs denoising on the input image. The ways of image denoising include but are not limited to techniques such as mean filtering, median filtering, and Gaussian filtering. These methods can effectively remove the noise in the image and ensure the clarity of the image.
[0020] In some embodiments, the image input module can also support the cropping operation of the screen image and the control image. The purpose of cropping is to focus on the control area and remove the unnecessary background part to facilitate subsequent image processing and control recognition. For a specific control area, the cropping function can shrink the image to the most relevant part, making the image more concentrated and accurate in the subsequent processing stage.
[0021] In this embodiment, the image input module highly collaborates with subsequent information geometric matching module, manifold learning module, Kalman filtering module, etc. Through the standardized image data provided by the image input module, the subsequent modules can achieve control positioning and automated test operations through means such as image matching, feature extraction, and dynamic prediction. During the connection process between the image input module and these subsequent modules, the standardization and processing quality of the image data directly affect the performance of the entire system and the accuracy of the test results.
[0022] Specifically, the image input module can support two main image input methods: one is to capture the entire application interface through a screen capture, and the other is to manually upload a control image by the user. To ensure the consistency of the image data, the image input module will automatically unify the format after receiving these images and convert the images into the format required by the subsequent processing module.
[0023] In certain embodiments, the image input module provides a user interface for the user to upload a control image. These images are usually screenshots of APP interface elements and can be obtained by manual or automatic capture. For the manual screenshot method, the user can directly select the control on the APP interface and take a screenshot through the function provided by the image input module.
[0024] As an option, the image input module can also be configured to automatically identify and extract the control images in the APP interface. The realization of this function depends on image processing algorithms, especially the control recognition technology based on template matching. In a possible implementation, the image input module automatically detects the control area in the image using the visual features of the control. This process greatly improves the efficiency of testing.
[0025] In some embodiments, the image input module is used in combination with the manifold learning module to improve the efficiency and accuracy of image processing. The principle of manifold learning is based on the low-dimensional embedding of image data. After the image input module inputs the image data into the system, the manifold learning module reduces the dimension of the image data through an algorithm and extracts effective geometric features. This process can be described by the following mathematical formula: x i =∑ j∈N(i) w ij x j ; where x i is the high-dimensional vector of the image, N(i) represents the neighboring points in the image, w ij is the contribution coefficient of the neighboring point to x i , and x j is the image feature of the neighboring point. After the image input module performs data input, the image data is sent to the manifold learning module for processing, and finally the image data is converted into a low-dimensional feature vector.
[0026] An image preprocessing module, configured to perform denoising, cropping, and contrast enhancement processing on the received image; In this embodiment, the image data received by the image preprocessing module comes from the image input module, which provides control images and screen images input through manual screenshot or a preset image library. These image data usually include a large amount of redundant information, and the image quality is sometimes affected by noise, blurring, etc. Therefore, it is necessary to clean and standardize these images through the image preprocessing module to ensure the normal operation of subsequent modules.
[0027] In some embodiments, the image preprocessing module first performs denoising processing. Noise is a common interference factor in image processing, which affects the image quality and makes it difficult to extract image features. Denoising methods can include techniques such as mean filtering, median filtering, and Gaussian filtering. Through these denoising algorithms, the noise part of the image will be effectively removed, and the image quality will be significantly improved.
[0028] Specifically, in a possible implementation, the image preprocessing module uses Gaussian filtering to denoise the input control image and screen image. Gaussian filtering blurs the image by applying a Gaussian function, thereby eliminating random noise in the image. Mathematically, the input image I′(x,y) is convolved with the Gaussian kernel G(i,j) to obtain the denoised image I ′ : where I′(x,y) represents the pixel value of the image after Gaussian filtering denoising at position (x,y); I(x+i,y+j) generally represents the pixel value at the coordinate (x,y) in the image I, and the pixel value at the new position after offset (i,j); G(i,j) is the weight value of the Gaussian kernel, usually a two-dimensional matrix, indicating the weighting degree of each neighboring pixel value in the convolution process; i and j respectively represent the offsets of the elements in the Gaussian kernel, ranging from -k to k, defining the window size of the convolution operation; x+i,y+j represents the pixel position after displacing the image by the offsets i and j.
[0029] Generally, the image preprocessing module also performs a cropping operation on the image. The purpose of cropping is to remove unnecessary background parts in the image, focus on the control area, and reduce possible interference factors in subsequent processing. Specifically, the control areas in the control image and screen image can be determined by edge detection or template matching methods, and then these areas are cropped out from the entire image. The cropped image only retains the parts related to the control, removing the redundant background area, which can effectively improve the accuracy of subsequent image matching.
[0030] As an option, the image preprocessing module can also further enhance the visual effect of the image by enhancing the contrast of the image. The enhancement of image contrast can be achieved through histogram equalization technology. The basic idea of histogram equalization is to adjust the distribution of image pixel values to improve the contrast of the image. Its mathematical formula is: where h(i) represents the frequency of pixel value i in the image; h ′ (i) represents the frequency after equalization; min(h) represents the minimum frequency in the frequency distribution; max(h) represents the maximum frequency in the frequency distribution.
[0031] In the image preprocessing module, the image can also be scaled, especially when the resolution of the input image is low. By uniformly scaling the image to a certain size, image matching problems caused by resolution differences can be avoided. In some embodiments, the image preprocessing module can automatically adjust the image size according to the resolution of the input image to facilitate subsequent module processing.
[0032] In some embodiments, the image preprocessing module performs color space conversion on the image. For example, the image is converted from the RGB color space to a grayscale image or an HSV color space. Grayscale images can better eliminate color interference and focus on the shape and structure of the control, while the HSV color space can extract more image features by separating hue, saturation, and brightness information. The mathematical formula for color space conversion is as follows: Y = 0.299R + 0.587G + 0.114B; Among them, R, G, B represent the red, green and blue components in the image respectively, and Y is the grayscale value after conversion; through the conversion, the color information of the image becomes more concise, which is conducive to subsequent feature extraction and matching.
[0033] Specifically, in the image preprocessing module, each image data from the image input module will go through the above multiple processing steps to ensure the best image quality. All preprocessed images will be passed to the manifold learning module for dimensionality reduction and feature extraction. The close cooperation between the image preprocessing module and the manifold learning module ensures the accuracy and effectiveness of the image data, and provides a solid foundation for subsequent image matching and control positioning.
[0034] In addition, the image preprocessing module's processing of input images is not limited to basic operations such as denoising, cropping, and enhancement, but also includes image formatting. Image formatting includes but is not limited to converting images to the standard size and file format required by the system, which can make the image more consistent in subsequent processing stages and avoid errors caused by inconsistent image formats.
[0035] Manifold learning module, used to reduce the dimension of the preprocessed image, extract the geometric features of the image, and embed the features of the control and background image in the low-dimensional manifold space; In this embodiment, the manifold learning module is closely connected with the aforementioned image preprocessing module and the subsequent image matching module. In the image preprocessing module, the image is processed by denoising, cropping, enhancement, etc. to obtain a high-quality standardized image. Then, these image data will be passed to the manifold learning module for dimensionality reduction and feature extraction. The output of the manifold learning module is a low-dimensional feature vector, which contains the core information of the image and can effectively support subsequent image matching and control positioning.
[0036] Manifold learning is a non-linear dimensionality reduction technique aimed at revealing the intrinsic structure in high-dimensional data by constructing a low-dimensional manifold. In some embodiments, the manifold learning module employs algorithms such as Locally Linear Embedding (LLE) or t-SNE to map the feature space of image data to a low-dimensional manifold space by mapping. Through these techniques, the manifold learning module can extract the core geometric features in the image, remove redundant information, and improve the performance of subsequent image processing modules.
[0037] Generally, the manifold learning module first constructs distance metrics and neighborhood relationships for the input high-dimensional image data, and then embeds the data into a low-dimensional space through the locally linear embedding algorithm. This process can effectively retain the important features in the image and remove unnecessary noise.
[0038] In this embodiment, the manifold learning module reduces the dimensionality of the image data by the Locally Linear Embedding (LLE) method. Specifically, the LLE algorithm processes the image data through the following steps: Construct a neighborhood graph: For each sample in the input image data, calculate its distance from other samples and select the nearest neighbor samples to construct a neighborhood graph. The neighborhood graph reflects the similarity between samples.
[0039] Assume that the input image data is a set of points X = {x1, x2,..., x N} in a high-dimensional space, where each x i is a high-dimensional data point. The manifold learning module first calculates the similarity or distance between data points, and common measurement methods include Euclidean distance or Manhattan distance.
[0040] Local reconstruction: For each data point x i , reconstruct its local representation through its neighborhood data points {x j |j ∈ N(i)}. This reconstruction process is carried out by minimizing the following reconstruction error: where x i is the i-th point in the original high-dimensional image data; N(i) is the neighborhood set adjacent to the point x i ; w ij is the weighting coefficient, indicating the contribution of the point x j to the reconstruction of the point x i ; ∥∥x i - ∑ j∈N(i) w ij x j ∥∥ 2 is the sum of the squares of the reconstruction errors for each data point x i .
[0041] Global embedding: By solving an optimization problem, an embedding representation Y = {y1, y2,..., y N} in a low-dimensional space is found such that the point set Y in the low-dimensional space maintains a geometric structure similar to the point set X in the high-dimensional space. Specifically, the optimization problem can be expressed as: where: y i is the embedding point in the low-dimensional space corresponding to the high-dimensional data point x i ; N(i) is the set of points adjacent to y i in the low-dimensional space; w ij is the reconstruction contribution of the embedding point y j to y i ; ∥∥y i - ∑ j∈N(i) w ij y j ∥∥ 2 is the sum of squared reconstruction errors in the low-dimensional embedding space.
[0042] In some embodiments, the manifold learning module reduces the dimension through the locally linear embedding algorithm and extracts the low-dimensional feature vectors of the image data. These feature vectors can not only reflect the overall structure of the image but also retain the local information in the image. For example, the shape, edge information, and texture features of the control can all be fully reflected in the low-dimensional space.
[0043] As an option, the manifold learning module can also combine other dimensionality reduction techniques, such as principal component analysis (PCA) or independent component analysis (ICA), to further improve the effect of feature extraction. PCA is a dimensionality reduction method based on linear transformation. It calculates the covariance matrix of the data and sorts its eigenvalues to select the most important principal components for dimensionality reduction. ICA further reduces redundant information by finding independent components that are uncorrelated in the data and extracts more discriminative features.
[0044] The output of the manifold learning module is the low-dimensional feature vectors, which will be passed as input data to the subsequent image matching module for image recognition and control positioning. Through the dimensionality reduction and feature extraction of the manifold learning module, the subsequent modules can perform control recognition and testing operations more efficiently, significantly improving the processing speed and accuracy of the system.
[0045] In some embodiments, the manifold learning module can also flexibly adjust the dimension of dimensionality reduction according to actual needs to adapt to the characteristics of different image data. If the image data contains relatively complex structural information, the manifold learning module can appropriately increase the dimension after dimensionality reduction to retain more detailed features and ensure the reliability of subsequent processing steps.
[0046] An information geometric matching module, which is used to calculate the similarity between the control image and the background image according to the geometric features of the image, and optimize the image matching accuracy based on mutual information and KL divergence; In this embodiment, the information geometric matching module depends on the low-dimensional feature vectors output by the manifold learning module as input. After manifold learning, these feature vectors contain the core geometric features of the image. The information geometric matching module will calculate the geometric distance to match the similarity between different images, so as to locate the corresponding control elements in the image.
[0047] The core idea of the information geometric matching module is to use the "distance" metric in information geometry theory to achieve image matching. Generally, the image matching problem can be transformed into the process of finding similar points in the feature space. Specifically, the similarity between two images is measured by calculating the distance between low-dimensional feature vectors.
[0048] Generally, the information geometric matching module performs image matching by calculating metrics such as Riemannian distance or Manhattan distance between feature vectors. Riemannian distance is widely used in manifold learning, and it can measure the geometric differences between points in the low-dimensional embedding space, thus helping to achieve accurate matching.
[0049] In this embodiment, the information geometric matching module uses Riemannian distance as the distance metric to measure the similarity of low-dimensional feature vectors. Specifically, assume that the feature vectors after being processed by the manifold learning module are y i and y j , then the Riemannian distance d(y i , y j ) between them can be calculated by the following formula: where: y i , y j are two feature vectors in the low-dimensional feature space respectively; G is the metric tensor, which represents the geometric structure of the image data and is usually learned according to the training data; (y i - y j ) T G(y i - y j ) is the weighted difference between the feature vectors, which reflects the geometric distance between two points.
[0050] In addition, the information geometric matching module can also use metrics such as Euclidean distance and cosine similarity for matching. The specific selection depends on the different requirements of the application scenario and image features. In a possible implementation, the information geometric matching module flexibly selects the matching algorithm according to different image features and control requirements to ensure higher matching accuracy.
[0051] During the image matching process, the information geometric matching module locates the controls precisely through an optimization algorithm. In some embodiments, the information geometric matching module adopts the method of minimizing the error function and achieves the optimal matching by continuously adjusting the control position. The definition of the error function can be: where: E is the error function, representing the matching error between the image feature vector and the target control feature vector; N is the number of image feature vectors (i.e., the number of feature points in the image); y i is the i-th feature vector in the image; y target is the feature vector of the target control; d(y i , y targct ) is the distance metric between y i and y target , which can be Riemannian distance, Euclidean distance or other suitable metrics.
[0052] The output of the information geometric matching module will be passed to the subsequent automated test module as the control positioning result for further test operations. Through the precise control position provided by the information geometric matching module, the automated test system can accurately perform interaction operations on the control, such as clicking, dragging, etc.
[0053] An image transformation module, which is used to perform diffusion processing on the image based on partial differential equations and perform affine transformation according to the changes of the image to ensure the geometric stability of the image; In this embodiment, the image transformation module is closely connected with the aforementioned image preprocessing module, manifold learning module and information geometric matching module. The image preprocessing module ensures the image quality through operations such as denoising and enhancement; the manifold learning module extracts image features through dimensionality reduction; the information geometric matching module calculates the geometric distance between images and performs matching. After obtaining the matching result, the image transformation module will perform geometric transformation operations on the image as needed, so that the control can be accurately located in different image spaces.
[0054] The image transformation module mainly relies on geometric transformation theories, including transformation methods such as translation, rotation, and scaling. Generally, the image transformation module calculates the required transformation parameters and performs the transformation according to the initial position of the control in the image and the expected position of the control in the target image. The core purpose of these transformation operations is to ensure that the control can be correctly recognized under different coordinate systems or perspectives.
[0055] Specifically, the image transformation module adjusts the image through methods such as affine transformation and perspective transformation. Affine transformation is a linear transformation that preserves parallelism and transforms the image through operations such as translation, rotation, and scaling; perspective transformation is applicable to more complex scenarios and can simulate the perspective effect in reality.
[0056] In this embodiment, the image transformation module first calculates the preliminary position of the control in the image and determines the transformation parameters according to the desired position in the target image or coordinate system. Assume that the position of the control in the image is (x ′ , y ′ ), and the target position is (x t , y t ), then the transformation of the image can be represented as a combination of operations such as translation, rotation, and scaling.
[0057] The formula for translation transformation is as follows: Where: (x, y) is the position of the control in the original image; (x ′ , y ′ ) is the position of the control after transformation; Δx and Δy respectively represent the horizontal and vertical offsets of translation.
[0058] The rotation transformation can be represented by the following matrix: Where θ is the rotation angle, (x ′ , y ′ ) is the position of the control after rotation, and (x, y) is the position of the control in the original image.
[0059] For the scaling transformation, its formula is: Where (x ′ , y ′ ) is the position of the control after rotation, (x, y) is the position of the control in the original image, and s is the scaling factor, indicating the increase or decrease in the size of the control.
[0060] In some embodiments, the image transformation module can also perform composite transformation by combining multiple transformation operations. For example, by combining translation, rotation, and scaling operations, complex control position adjustment can be achieved. The transformed image will be sent to the subsequent image matching and control positioning module to ensure the precise positioning of the control in different image spaces.
[0061] The output of the image transformation module will be used as the input of the adjusted image and passed to the subsequent image matching module and automated testing module. In the image matching module, the transformed image will be matched with the target image to ensure that the controls can be accurately located in different images. In the automated testing module, the positions of the controls after image transformation will be used to perform various interaction operations, such as clicking and dragging, to ensure the smooth progress of automated testing.
[0062] A dynamic element processing module for dynamically predicting and correcting the positions of controls during the image matching process through Kalman filtering; In this embodiment, the dynamic element processing module is closely connected to the aforementioned image transformation module and information geometric matching module. After the image transformation module adjusts the positions of the controls by methods such as translation, rotation, and scaling, the dynamic element processing module continues to track and process the dynamic elements in the image. By timely updating and re-identifying the dynamic elements, the robustness and stability of the system in response to interface changes are ensured.
[0063] One of the core tasks of the dynamic element processing module is the recognition and tracking of dynamic elements. Generally, the dynamic elements in the image may change in position or state due to user operations, animation effects, or other changes. To ensure the correct positioning of the controls, the dynamic element processing module must be able to recognize these changes and adjust the positioning strategy in real time.
[0064] In this embodiment, the recognition process of dynamic elements is carried out through change detection based on image features. By comparing consecutive image frames, the module can detect changes in the controls, such as position offsets or appearance changes. To improve the accuracy of dynamic element recognition, the dynamic element processing module can use techniques such as optical flow method and Kalman filtering to predict and compensate the movement trajectories of dynamic elements.
[0065] The optical flow method calculates the movement direction and speed of each pixel by analyzing the movement of pixels in the image. This is particularly effective for detecting and tracking dynamic controls. The formula is expressed as: I x u + I y v + I t =0; Where: I x and I y are the gradients of the image in the x and y directions respectively; u and v are the movement speeds of the pixels in the x and y directions respectively; I t is the change of the image over time.
[0066] Kalman filtering uses a dynamic model to predict and correct the position of the target control. In the movement trajectory of dynamic elements, Kalman filtering can reduce uncertainty and noise by continuously updating the current position and movement state of the control, providing a more accurate estimate of the control position.
[0067] During the image transformation process, dynamic elements may be affected by operations such as translation, scaling, and rotation of the image. To ensure that the control can still be accurately recognized after transformation, the dynamic element processing module must adapt to these changes.
[0068] Specifically, the dynamic elements after image transformation can be repositioned by calculating their initial positions and transformation parameters. For example, when the control position is translated, the module adjusts the position information of the dynamic element by recording the translation offset; when the control rotates, the module updates the coordinates of the control through a rotation matrix. The following is the translation transformation formula: Where: (x,y) is the position of the control in the original image; (x ′ ,y ′ ) is the position of the control after transformation; Δx and Δy represent the horizontal and vertical offsets of the translation respectively.
[0069] The output of the dynamic element processing module will directly affect the operations of the subsequent control positioning module and the automated testing module. In the image matching module, the image processed by the dynamic element will ensure that the control can be accurately positioned in the changing UI interface. In the automated testing module, the real-time control position provided by the dynamic element processing module will be used to perform relevant interaction operations to ensure that the system can correctly simulate user operations under the dynamic interface.
[0070] The automated operation execution module is used to execute corresponding automated operations after the control is successfully positioned; In this embodiment, the automated operation execution module closely cooperates with the foregoing image transformation module, dynamic element processing module, information geometric matching module, etc. The image transformation module provides the position of the control after transformation, the dynamic element processing module updates the control position in real time according to the dynamic changes, and the information geometric matching module accurately matches the position of the control. The automated operation execution module then accurately performs operations related to the control based on this information to ensure the smooth progress of the automated testing.
[0071] The core principle of the automated operation execution module is to simulate the behavior of human users through the input of accurate control positions and operation types to interact with the control. Generally, these operations include actions such as clicking, inputting, swiping, and dragging.
[0072] Specifically, the automated operation execution module simulates these operations using input devices such as the mouse and keyboard by receiving the position coordinates and operation commands of the control. For click operations, the module determines the coordinates (x, y) of the control, moves the mouse to that position, and simulates a mouse click event; for input operations, the module inputs the entered content by simulating keyboard input; for swipe or drag operations, the module drags the mouse according to the position of the control.
[0073] In this embodiment, the automated operation execution module first calculates the exact position of the control based on the control position information obtained from the information geometric matching module. Assuming the target position of the control is (x, y), the automated operation execution module performs the operation through the following steps: Click operation: Based on the control position (x, y), the automated operation execution module moves the mouse pointer to that position and simulates a mouse click event. The implementation of the click operation can be achieved by calling the mouse click function: Click(x,y)where(x,y)is the target position; Where: x is the coordinate of the target position in the horizontal direction; y is the coordinate of the target position in the vertical direction; (x, y) is the position of the target control on the screen or application interface; Click(x, y) represents simulating a mouse click operation at the coordinates (x, y).
[0074] Input operation: For controls such as text input boxes, the module inputs the text content into the target control by simulating keyboard input. Assuming the entered text is t, the process of performing the input operation is: Input(t)into(x,y); Where: t is the text content to be entered, which can be any string, such as letters, numbers, symbols, etc.; (x, y) is the position of the target control, representing the coordinates of the input box or text box in the interface; Input(t)into(x, y) represents inputting the text t into the control corresponding to the position (x, y).
[0075] Swipe operation: For controls that need to be swiped (such as scroll bars), the automated operation execution module simulates mouse scroll or drag events to ensure that the control can be swiped as required. The swipe operation can be achieved in the following way: Scroll(x1,y1,x2,y2)where(x1,y1)and(x2,y2)are start and end positions; Among them, (x1, y1) is the starting position of the swipe, usually a coordinate point on the control or page; (x2, y2) is the ending position of the swipe, indicating the position of the control or page after the swipe ends; Scroll(x1, y1, x2, y2) represents performing a swipe operation from the starting position (x1, y1) to the ending position (x2, y2).
[0076] The automated operation execution module not only performs operations but also can adjust the operation process according to the feedback information. In some embodiments, the module will confirm whether the operation is successful by monitoring the real-time state of the control. For example, when clicking a button, the module will monitor the response state of the button to ensure that the button is clicked correctly; when inputting text, the module will verify whether the content in the input box is filled in correctly. If the operation fails, the module will retry or report an error according to the set strategy.
[0077] The output of the automated operation execution module will be part of the subsequent test results. By docking with the subsequent verification module, the entire automated test process is completed. Through the automated operations performed by this module, the system can simulate real user interaction processes, thereby verifying the functions, stability, and performance of the application.
[0078] The report generation module is used to record and generate the operation results and test reports of the test process; In this embodiment, the report generation module is closely connected to the foregoing various modules (such as the image input module, image preprocessing module, image transformation module, automated operation execution module, etc.). The automated operation execution module will generate an operation log when performing an operation. The image transformation module and the control positioning module will record the position information of the control. The dynamic element processing module will update the state changes of the control in real time. All these information will be summarized and transmitted to the report generation module, and this module will complete the recording of the test results and the generation of the report.
[0079] One of the core tasks of the report generation module is to collect key information during test execution from various modules and effectively analyze this data. In some embodiments, the module will collect the following types of data: Operation log: Records various types of user operations simulated by the automated operation execution module during the execution process, including clicks, inputs, swipes, etc.; Control state: Records the state changes of each control during the test, such as position changes, content changes, response states, etc.; Test results: Records the execution results of each test case, including information such as pass, fail, warning, etc.; Performance metrics: Include performance data such as page load time, response time, etc.
[0080] During the data analysis process, the report generation module can perform statistics, summarization, and analysis on this data to clearly present it to the testers. Specifically, the module will generate information such as the pass rate, failure rate, and warnings for each test case by comparing the expected results with the actual execution results, and display it in the form of charts or tables.
[0081] In this embodiment, the report generation module can automatically generate a formatted test report based on the collected test data. The report generation module can generate reports in different formats according to user requirements, such as: Text report: It details each test step, test result, and related logs, and is suitable for detailed review.
[0082] Table report: It lists information such as the execution status of each test case and the status changes of controls in tabular form.
[0083] Graphical report: It shows data such as performance metrics and pass rates during the test process through charts.
[0084] The generated report usually includes the following contents: Test overview: It outlines the test objectives, execution environment, and the number of test cases; Test execution process: It gradually records the execution process, input data, execution results, etc. of each test case; Test result analysis: It analyzes the test pass rate, failure rate, error messages, etc. based on the execution results; Performance analysis: It evaluates the performance of the system or application and provides data such as page load time and response time; Summary and suggestions: It gives improvement suggestions or subsequent test plans based on the test results.
[0085] The report generation module generates a test report format that meets the user's requirements according to the needs. The report can be in PDF format, Excel spreadsheet, or HTML page, etc., which is convenient for subsequent viewing, archiving, and sharing.
[0086] When generating the report, the report generation module may use some mathematical formulas to display performance data or test results. Taking the pass rate as an example, assuming the total number of test cases is N and the number of passed cases is P, then the pass rate R can be calculated by the following formula: Among them, R is the pass rate, P is the number of passed test cases, and N is the total number of test cases.
[0087] Similarly, the performance analysis report may include some common calculations, such as performance metrics like response time and load time. Assuming the response times of requests during the test are t1, t2,..., t n , then the average response time Tavg It can be calculated by the following formula: Where, T avg is the average response time, t i is the response time of the i-th test, and n is the number of tests.
[0088] The output of the report generation module will directly affect the tester's evaluation of the system performance, stability and functions. Through the generated test report, the tester can quickly understand the test results, locate the problems or performance bottlenecks existing in the system, and formulate subsequent test plans or system improvement plans according to the analysis results in the report. The report generation module can also support the timed task of automatically generating test reports to ensure the efficiency of the test process.
[0089] Please refer to Figure 5 , the automated testing method for the obfuscated APP based on visual recognition, including the following steps: S1. Receive the target control image and the screen image to be recognized; The main task of this step is to receive the target control image to be processed and the screen image to be recognized from the system or the user. The target control image is usually the standard or template image of the control, while the screen image is the current screenshot of the application interface containing the control. The input of this step is two images, which will be used as the basic data for subsequent processing.
[0090] S2. Preprocess the input images, including denoising, cropping, and contrast enhancement operations; In this stage, the input images will undergo a variety of preprocessing operations. The denoising process can remove the unnecessary noise in the images and improve the clarity of the images; the cropping operation is used to remove the irrelevant background information and focus on the control area; the contrast enhancement technology adjusts the brightness and contrast of the images to make the edges of the controls clearer, facilitating subsequent image matching and feature extraction.
[0091] S3. Use manifold learning to reduce the dimension of the preprocessed images and extract the low-dimensional geometric features of the images; Manifold learning is an effective non-linear dimensionality reduction method for extracting meaningful low-dimensional features from high-dimensional data. In this step, the preprocessed images are reduced in dimension through the manifold learning algorithm, and the low-dimensional geometric features of the images are extracted. These features will contain the geometric and shape information between the control image and the screen image, providing strong support for image matching.
[0092] S4. Calculate the similarity between the control image and the screen image based on information geometry, and optimize the matching accuracy by calculating the mutual information value and KL divergence; In this stage, the information geometry theory is used to calculate the similarity between the control image and the screen image. The mutual information value is used to measure the common information between two images, and the KL divergence is used to measure the difference between two images. These calculations help to optimize the matching accuracy and improve the accuracy of control positioning.
[0093] S5. Diffuse the image through partial differential equations and apply affine transformation to correct the geometric changes of the image; Partial differential equations (PDEs) can be used for the diffusion process in image processing to help smooth the image and reduce noise. In this way, the detailed information of the image can be effectively preserved and the background noise can be suppressed. At the same time, affine transformation is used to correct the geometric deformations in the image, including rotation, scaling, etc., to ensure the spatial consistency of the image.
[0094] S6. Dynamically predict and correct the control position through Kalman filtering; Kalman filtering is a recursive filtering method that can predict the control position based on the previous state and current observations of the control. Through this process, Kalman filtering can update the position estimate of the control in real time and correct the errors caused by image processing or environmental changes, thereby improving the accuracy and stability of control positioning.
[0095] S7. After successful control positioning, perform automated test operations; After completing the control positioning, the system will perform automated test operations, such as simulating clicks, inputs, or swipes. Through accurate control positioning, the system can perform corresponding operations at the correct positions, thereby realizing the automated execution of automated test tasks.
[0096] S8. Record the test results and generate a test report; Finally, the system will record the results of the entire automated test process, including key information such as whether the operations are successful, execution time, and test pass rate. This information will be summarized and a detailed test report will be generated to help testers evaluate the test results, discover potential problems, and provide improvement suggestions. The test report can include data in the form of text, tables, or graphs and can be exported as report files in different formats according to requirements.
[0097] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An automated testing system for obfuscated APPs based on visual recognition, characterized in that, Comprising: An image input module for receiving a target control image and an APP screen image; An image preprocessing module for performing denoising, cropping, and contrast enhancement on the received images; A manifold learning module for performing dimensionality reduction on the preprocessed images, extracting geometric features of the images, and performing feature embedding on the control image and the background image in a low-dimensional manifold space; An information geometric matching module for calculating the similarity between the control image and the background image according to the geometric features of the images, and optimizing the image matching accuracy based on mutual information and KL divergence; An image transformation module for performing diffusion processing on the images based on partial differential equations and performing affine transformation according to the changes in the images to ensure the geometric stability of the images; A dynamic element processing module for dynamically predicting and correcting the position of the control during the image matching process through Kalman filtering; An automated operation execution module for performing corresponding automated operations after successful control positioning; A report generation module for recording and generating the operation results and test reports of the test process.
2. The automated testing system for the obfuscated APP based on visual recognition according to claim 1, characterized in that, The image input module is used to receive an image library of manual screenshots from the user and input the control image and the screen image into the system.
3. The automated testing system for the obfuscated APP based on visual recognition according to claim 1, wherein The manifold learning module includes a locally linear embedding unit for performing dimensionality reduction on the image data to generate low-dimensional feature vectors.
4. The automated testing system for the obfuscated APP based on visual recognition according to claim 1, wherein The information geometric matching module includes a mutual information calculation unit and a KL divergence calculation unit. The mutual information calculation unit is used to calculate the mutual information value according to the joint probability distribution of the control image and the background image, and the KL divergence calculation unit is used to calculate the KL divergence between the control image and the background image and optimize the matching result according to the mutual information and the KL divergence.
5. The automated testing system for the obfuscated APP based on visual recognition according to claim 1, characterized in that The image transformation module is used to execute a heat diffusion equation on the control image and perform geometric correction of the image by applying an affine transformation matrix according to the image deformation.
6. The automated testing system for the obfuscated APP based on visual recognition according to claim 1, wherein The dynamic element processing module includes a Kalman filtering unit for predicting and correcting the error of the control position according to the historical image matching results and the current observation data.
7. The automated testing system for the obfuscated APP based on visual recognition according to claim 6, wherein The state prediction equation of the Kalman filtering unit is: Among them, represents the predicted state at time t; A is the state transition matrix, which describes the state evolution of the system from time t - 1 to time t, and it represents the dynamics of the system; represents the state estimate at time t - 1, that is, the predicted state at the previous moment; B is the control input matrix; u t is the control vector, representing the control input at time t.
8. The automated testing system for the obfuscated APP based on visual recognition according to claim 6, characterized in that, The state update equation of the Kalman filtering unit is: Kalman gain calculation: K t = P t H T (HP t H T + R) -1 ; where K t is the Kalman gain matrix, which determines the weighted ratio between the predicted value and the actual observed value; P t is the prediction error covariance matrix, representing the uncertainty of the predicted state; H T represents the transpose of matrix H; H is the observation matrix, which describes the transformation from the state space to the observation space; R is the observation noise covariance matrix, representing the noise in the observation process; State update: where x t is the updated state estimate, i.e., the state adjusted according to the observation information; is the predicted state estimate, representing the current state obtained from the prediction at the previous moment; K t represents the Kalman gain matrix; z t is the actual observation value, the true measurement of the system; is to convert the predicted state into the observation space according to the observation matrix H; Error covariance update: P t = (I - K t H)P t ; Among them, P t is the updated error covariance matrix, representing the updated state uncertainty; I is the identity matrix, with the dimension consistent with the state space; K t H is the product of the Kalman gain and the observation matrix, reflecting the impact of measurement update on uncertainty.
9. The automated test system for the obfuscated APP based on visual recognition according to claim 1, wherein, The automated operation execution module is used to perform click and swipe interaction operations by the control operation unit after the matching accuracy between the control image and the screen image calculated by the information geometric matching module reaches a preset threshold, and when the matching accuracy calculated by the information geometric matching module is lower than the preset threshold, the automated operation execution module automatically performs a retry operation, with a maximum of three retries.
10. The automated testing method for the obfuscated APP based on visual recognition is applied to the automated testing system for the obfuscated APP based on visual recognition according to any one of claims 1-9, and is characterized in that, Including the following steps: S1. Receive a target control image and a screen image to be recognized; S2. Preprocess the input images, including denoising, cropping, and contrast enhancement operations; S3. Use manifold learning to perform dimensionality reduction on the preprocessed images and extract low-dimensional geometric features of the images; S4. Calculate the similarity between the control image and the screen image based on information geometry, and optimize the matching accuracy by calculating the mutual information value and the KL divergence; S5. Perform diffusion processing on the image through partial differential equations and apply affine transformation to correct the geometric changes of the image; S6. Dynamically predict and correct the error of the control position through Kalman filtering; S7. After the control is successfully positioned, perform automated test operations; S8. Record the test results and generate a test report.
Citation Information
Cited By
AI visual positioning method and system for robot automatic assembly
CN120839797A