Cross-platform GUI (Graphical User Interface) test script generation method based on multi-modal visual identification

Through multimodal visual recognition technology, the YOLOv5 and CRNN-OCR model combined with CNN to generate cross-platform GUI test scripts is solved, which solves the selector failure problem caused by front-end framework upgrades and improves the accuracy and efficiency of the test.

CN120234252AActive Publication Date: 2025-07-01四川互慧软件有限公司
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510427182.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-01
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

In the prior art, the GUI testing tool has high selector failure efficiency and insufficient visual adaptation when the front-end framework is upgraded, resulting in large positioning deviations, making it difficult to achieve cross-platform intelligent script generation.

Method used

The multimodal visual recognition method is adopted, element detection and CRNN-OCR model are used for text recognition through the YOLOv5 model, feature fusion and layout relationship analysis are combined with CNN to generate cross-platform GUI test scripts.

Benefits of technology

Reduces the probability of selector failure due to front-end framework upgrades, improves the accuracy and efficiency of cross-platform testing, and simplifies workflow.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234252A_ABST
    Figure CN120234252A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of software test automation, in particular to a cross-platform GUI (Graphical User Interface) test script generation method based on multi-modal visual identification, which mainly comprises the following steps of: respectively sending preprocessed basic data to a visual channel and a text channel for processing; the visual channel performs element detection on the basic data through a YOLOv5 model to obtain a first recognition result, and the text channel performs character recognition on the basic data through a CRNN-OCR model to obtain a second recognition result; and carrying out feature fusion on the first recognition result and the second recognition result, and outputting to a script generator. According to the method, a conventional tool is not used for processing any more, a dual-channel processing layer is adopted for analyzing and extracting the current basic data, the working process is simplified, and the probability that a large number of selectors fail due to upgrading of a front-end framework is greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of software test automation, and more specifically, to a cross-platform GUI test script generation method based on multi-modal visual recognition. Background Art

[0002] GUI stands for Graphical User Interface. In the prior art, event streams are usually generated by randomly interacting with screen coordinates, and tests are guided by constructing interface state machines or calling graphs, relying on predefined rules or scripts to control interface elements. And based on the natural language understanding and generation capabilities of LLM, test requirements can be directly converted into operation sequences, and the parameter dependencies between interfaces can be analyzed to achieve intelligent script generation for multi-step scenarios.

[0003] However, traditional tools (such as Selenium) rely on HTML structure parsing, and the upgrade of front-end frameworks causes many selectors to become invalid; insufficient visual adaptation: Recording tools based on coordinate positioning (such as QTP) have large positioning deviations when the resolution changes. Summary of the Invention

[0004] The purpose of the present invention is to provide a cross-platform GUI test script generation method based on multi-modal visual recognition to solve the above problems in the prior art.

[0005] The present invention is achieved through the following technical solutions: A cross-platform GUI test script generation method based on multi-modal visual recognition, comprising: Obtain the basic data to be processed, where the basic data includes image data or video data, and preprocess the basic data; Send the preprocessed basic data to the visual channel and the text channel for processing respectively. The visual channel performs element detection on the basic data through the YOLOv5 model to obtain the first recognition result, and the text channel performs character recognition on the basic data through the CRNN-OCR model to obtain the second recognition result; Fuse the features of the first recognition result and the second recognition result, perform layout relationship parsing through CNN, output the obtained parsing result to the script generator, and generate executable test scripts respectively through the script generator.

[0006] Preferably, the preprocessing of the basic data includes: Normalize the resolution of the basic data and output the first basic data, eliminate noise from the first basic data, and then perform dynamic element filtering to output the second basic data; Perform image enhancement on the second basic data to output the processed basic data.

[0007] Preferably, the dynamic element filtering includes: Set static conditions and dynamic conditions. The setting of static conditions includes setting a filtering threshold based on a numerical range or string matching, and the setting of dynamic conditions includes obtaining input or real-time data to generate conditions; Traverse the data set and set a condition judgment model. The condition judgment model includes making a logical judgment on the attributes of elements to confirm whether the attributes of elements are within the set filtering threshold, and outputting a data set that meets the conditions after the judgment.

[0008] Preferably, the feature fusion of the first recognition result and the second recognition result includes cross-modal attention alignment, including: Perform bilinear attention interaction:

[0009] Perform hierarchical attention aggregation, including using a multi-head mechanism to calculate multi-granularity associations in parallel:

[0010] Perform residual feature fusion:

[0011]

[0012] Set a loss function and perform contrastive loss optimization; In the formula, is the cross-modal similarity score between the first recognition result and the second recognition result, is the device of the first recognition result, is the th word or embedding vector in the text sequence of the second recognition result, is the multi-head attention function, is the vector concatenation function, is the calculation result of the first to the kth attention heads, is the linear transformation matrix of the multi-head attention output, is the first recognition result after performing residual feature fusion, is the initial first recognition result, is the layer normalization function, is the association strength between each row corresponding to a visual region and all text words, is the second recognition result, and are both projection matrices, is the second recognition result after performing residual feature fusion, is transpose of.

[0013] Preferably, the feature fusion includes: Performing vector normalization on the first recognition result and the second recognition result; Concatenating the first recognition result and the second recognition result into a joint vector:

[0014] wherein, is the joint vector obtained by fusion, is the calculated weight.

[0015] Preferably, the layout relationship parsing through CNN includes: Constructing a weighted spatial relationship matrix, inferring the hierarchical relationship through 3-layer GNN iteration, and outputting a DOM tree with semantic annotations, where the DOM tree includes a parent node and the first input box child node, the second input box child node, and the login button child node associated with the parent node.

[0016] Preferably, the construction of the weighted spatial relationship matrix includes:

[0017] wherein, is the element in the th row and th column of the spatial relationship matrix, is the natural exponential function, and are the coordinates of the th point, is the coordinate of the th point, is the bandwidth parameter of the Gaussian kernel function.

[0018] Preferably, the hierarchical relationship inference through 3-layer GNN iteration includes: Constructing the first layer of GNN, the second layer of GNN, and the third layer of GNN, where the first layer of GNN is used to process the original node features, the second layer of GNN is used to input the node embeddings of the first layer of GNN into a coarsening module, and the third layer of GNN aggregates the clusters of the second layer to form semantic units, and maps the clusters of the target layer back to the node relationship of the original graph through a full expansion function.

[0019] Preferably, the sending to the visual channel and text channel for processing includes: Judging whether the recognition is successful. If the first recognition result and the second recognition result are successfully recognized, feature fusion is performed; If the recognition is not successful, the detection threshold is adjusted, SIFT feature matching is adopted, and it is judged again whether the recognition is successful. If the recognition is successful, the current detection threshold is saved. If the recognition is not successful, an alarm signal is sent and an error report is output.

[0020] Preferably, the script generator generates executable test scripts respectively, including: Operating an inference finite state machine, including determining that the element type of the currently input is a login button or an input box. If it is a login button, adjacent text is checked to generate a click operation script. If it is an input box, a label associated with the input box is found to generate a type operation script; Writing an assertion generation algorithm, where the assertions include text assertions, structural assertions, and visual assertions.

[0021] The technical solution of the present invention has at least the following advantages and beneficial effects: Through the method provided by the present invention, it mainly includes sending the preprocessed basic data to the visual channel and the text channel for processing respectively. The visual channel performs element detection on the basic data through the YOLOv5 model to obtain a first recognition result, and the text channel performs character recognition on the basic data through the CRNN-OCR model to obtain a second recognition result; the first recognition result and the second recognition result are feature-fused and output to the script generator. Through the above method, conventional traditional tools are no longer used for processing, but a two-channel processing layer is adopted to analyze and extract the current basic data respectively, simplifying the work process and greatly reducing the probability that a large number of selectors become invalid due to the upgrade of the front-end framework. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0023] Figure 1 is the overall control flow chart of the present invention; Figure 2 is the schematic diagram of the processing flow of the visual channel and the text channel of the present invention.

[0024] Figure 3 is the schematic diagram of the processing flow of the layout relationship parsing of the present invention; Figure 4 is the schematic diagram of the processing flow of the element overlap conflict of the present invention; Figure 5 is the schematic diagram of the processing flow of the operation inference finite state machine of the present invention; Figure 6 is the schematic diagram of the processing flow of the element recognition failure of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Components of the embodiments of the present invention described and illustrated herein can be arranged and designed in a variety of different configurations.

[0026] The division of modules in this application is a logical division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.

[0027] The independently described modules or sub-modules can be physically separated or not: they can be implemented in software or in hardware, and some of the modules or sub-modules can be implemented in software, and the functions of these modules or sub-modules are called by the processor to implement this part of the module or sub-module, and other parts of the module or sub-module are implemented in hardware, for example, through a hardware circuit. In addition, some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this application.

[0028] Please refer to Figures 1-6 , the present invention provides a cross-platform GUI test script generation method based on multi-modal visual recognition, including: S101: Obtain the basic data to be processed, where the basic data includes image data or video data, and preprocess the basic data; More specifically about the preprocessing, the basic data is subjected to resolution standardization and then the first basic data is output. After noise elimination is performed on the first basic data, dynamic element filtering is performed to output the second basic data; image enhancement is performed on the second basic data to output the processed basic data.

[0029] Among them, image enhancement includes, after loading the image, usually grayscale conversion (such as converting to the HSV or YCrCb space using cv2.cvtColor), and noise reduction through filtering (mean filtering, median filtering). According to the target, an enhancement method is selected. For example, histogram equalization: calculate the grayscale histogram, and redistribute the pixel values through the cumulative distribution function, which is suitable for low-contrast images; Laplacian operator enhancement: highlight the edge structure, and noise reduction preprocessing needs to be combined to prevent noise amplification; local contrast enhancement (CLAHE): adjust the contrast in blocks to solve the problem of local detail loss caused by global equalization; adjust brightness / contrast (such as Gamma correction), or combine guided filtering to suppress noise.

[0030] S102: Send the preprocessed basic data to the visual channel and the text channel for processing respectively. The visual channel uses the YOLOv5 model to detect elements in the basic data to obtain the first recognition result, and the text channel uses the CRNN-OCR model to recognize text in the basic data to obtain the second recognition result; The obtained first recognition result or second recognition result is mainly a feature matrix or a feature vector, which is convenient for subsequent processing. YOLOv5 achieves a flexible balance between speed and accuracy through multiple model versions (YOLOv5s / m / l / x). Among them, YOLOv5s controls the network depth and width by adjusting the depth_multiple and width_multiple parameters, and the number of parameters is only 7.2M. The inference speed can reach 140FPS (Tesla V100), which is suitable for mobile deployment. Its improved CSP (CrossStage Partial) structure reduces redundant calculations, while the SPPF (Spatial Pyramid Pooling Fast) module changes the parallel max pooling to serial, improving the calculation efficiency by 30% while maintaining the multi-scale feature fusion ability. Secondly, before training, the optimal anchor boxes are dynamically generated through K-means clustering and genetic algorithms, reducing the cost of manual parameter tuning and improving the matching accuracy of detection boxes. At the same time, an attention mechanism is added: a CBAM module is added at the end of the Backbone to improve the small target detection ability.

[0031] The CRNN-OCR model does not require character segmentation. It directly extracts image features through CNN, models sequence dependencies through RNN, and transcribes and outputs text through CTC, realizing one-stop recognition from image to text. The CTC loss function allows the input and output sequence lengths to be misaligned, eliminating the need for character-by-character annotation and significantly reducing the data annotation cost. In addition, the improved Sauvola algorithm is used to dynamically calculate the window size and correct the candidate sequence based on the GUI domain dictionary.

[0032] S103: Fuse the features of the first recognition result and the second recognition result, parse the layout relationship through CNN, and output the obtained parsing result to the script generator to generate executable test scripts respectively through the script generator.

[0033] In this embodiment, the generated executable test scripts are mainly Playwright scripts, Puppeteer scripts, and HTML test reports, which are output by the output layer.

[0034] Through the method provided by the present invention, it mainly includes respectively sending the preprocessed basic data to the visual channel and the text channel for processing. The visual channel uses the YOLOv5 model to perform element detection on the basic data to obtain the first recognition result, and the text channel uses the CRNN-OCR model to perform character recognition on the basic data to obtain the second recognition result; the first recognition result and the second recognition result are subjected to feature fusion and output to the script generator. Through the above method, instead of using conventional traditional tools for processing, a two-channel processing layer is adopted to analyze and extract the current basic data respectively, which simplifies the work process and greatly reduces the probability of failure of many selectors due to the upgrade of the front-end framework.

[0035] An exemplary embodiment of the present invention, the dynamic element filtering includes: Set static conditions and dynamic conditions. The setting of static conditions includes setting a filtering threshold based on a numerical range or string matching. The setting of dynamic conditions includes obtaining input or real-time data to generate conditions; traverse the data set and set a condition judgment model. The condition judgment model includes making a logical judgment on the attributes of the elements to confirm whether the attributes of the elements are within the set filtering threshold, and after the judgment, output the data set that meets the conditions after filtering.

[0036] In the above steps, the algorithms involved include: Recursive filtering algorithm, check the child nodes layer by layer from the root node, and retain the parent node if the child nodes match; semantic segmentation and dynamic mask generation, combine an online instance segmentation model (such as Track Anything) to generate a semantic mask of dynamic objects, and optimize the accuracy through photometric consistency verification. Application scenario: distinguish static background and moving objects in dynamic SLAM to improve the robustness of mapping and positioning; range-based filtering function, implementation logic: traverse the object array and check whether the attributes fall within the specified range; higher-order functions and Lambda expressions, Python's filter(): supports custom functions or Lambda expressions, such as filtering even numbers; dynamic kernel parameter filtering (eBPF), implementation steps: parse the filtering expression to generate an AST, convert the AST to BPF instructions through BTF information, and dynamically replace the filtering logic in the BPF program.

[0037] In addition, after preprocessing the basic data, an anti-interference dynamic processing mechanism is also adopted, including temporal difference method (inter-frame change detection), feature matching (ORB feature point comparison), and semantic filtering (keyword blacklist mechanism).

[0038] Specifically, the temporal difference method includes inter-frame difference calculation, performing pixel-by-pixel grayscale difference on two or three consecutive frames of images; adaptive threshold segmentation, dynamically adjusting the threshold to improve robustness, for example, based on the statistical characteristics of the difference image; morphological optimization, using dilation and erosion operations to fill holes and remove noise.

[0039] Feature matching includes FAST key point detection, detecting whether 12 consecutive points among 16 pixels around a pixel point satisfy the brightness difference threshold T (such as 20% of the central brightness), and screening corner points; direction calculation (gray centroid method), calculating the moment of the key point neighborhood to determine the main direction; BRIEF descriptor generation, randomly selecting 256 pairs of pixel points in the key point neighborhood to generate a 256-bit binary vector; feature matching (Hamming distance) calculation of descriptor differences.

[0040] Semantic filtering includes text preprocessing, using tools (such as Jieba segmentation) to segment the text and filter out meaningless words (such as "de", "shi"); keyword matching, Boolean model: directly matching keywords in the blacklist; dynamic update mechanism, weighted updating of the blacklist based on user feedback.

[0041] In an exemplary embodiment of the present invention, the feature fusion of the first recognition result and the second recognition result includes cross-modal attention alignment, including: Performing bilinear attention interaction:

[0042] Performing hierarchical attention aggregation, including using a multi-head mechanism to calculate multi-granularity associations in parallel:

[0043] Performing residual feature fusion:

[0044]

[0045] Setting a loss function and performing contrastive loss optimization; In the formula, is the cross-modal similarity score between the first recognition result and the second recognition result, is the device of the first recognition result, is the th word or embedding vector in the text sequence of the second recognition result, is the multi-head attention function, is the vector concatenation function, is the calculation result from the 1st to the kth attention head, is the linear transformation matrix of the multi-head attention output, is the first recognition result after performing residual feature fusion, is the initial first recognition result, is the layer normalization function, is the association strength between each row corresponding to a visual area and all text words, is the second recognition result, and are both projection matrices, is the second recognition result after residual feature fusion, is the transpose of.

[0046] Map the text position to the visual element coordinate system through affine transformation. When there is a conflict between visual classification and text semantics (such as detecting a button but the text is "Cancel"), start the rule engine arbitration.

[0047] Specifically, the feature fusion includes: Perform vector normalization on the first recognition result and the second recognition result; Concatenate the first recognition result and the second recognition result into a joint vector:

[0048] In the formula, is the fused joint vector, is to calculate the weight.

[0049] An exemplary implementation of the present invention for parsing layout relationships through CNN includes: Construct a weighted spatial relationship matrix, infer the hierarchical relationship through 3-layer GNN iteration, and output a DOM tree with semantic annotations. The DOM tree includes a parent node and the first input box child node, the second input box child node, and the login button child node associated with the parent node.

[0050] Among them, if there is a conflict of element overlap, obtain the current conflict type, judge whether it is text coverage or area inclusion. If it is text coverage, calculate the z-index priority, and then adjust the element hierarchy. If it is area inclusion, establish a parent-child relationship, verify the semantic consistency. If they are consistent, merge them into a container node. If they are inconsistent, split them into independent nodes.

[0051] Among them, the priority calculation of the z-index involves the rules of the stacking context and the interaction of element attributes. The z-index attribute only takes effect when the position attribute of the element is relative, absolute, fixed, or sticky. The z-index value of an element with no positioning (position: static) is invalid. Among sibling elements, the element that comes later in the document flow covers the previous element by default. If the z-index is not set, the positioned element will cover the non-positioned element.

[0052] When both of two sibling elements are positioned and the z-index is not auto: the one with a larger z-index value covers the one with a smaller value (e.g., z-index: 2 covers z-index: 1); if the z-index values are the same, the element that comes later in the document flow covers the previous one.

[0053] In the same stacking context, elements are arranged from bottom to top in the following order (with increasing priority): Background and borders (background of the parent element); positioned elements with a negative z-index; block-level non-positioned elements; floating elements; inline non-positioned elements; positioned elements with z-index: 0 or auto; positioned elements with a positive z-index.

[0054] In an exemplary embodiment of the present invention, constructing a weighted spatial relationship matrix includes:

[0055] In the formula, is the element in the th row and th column of the spatial relationship matrix, is the natural exponential function, and are the coordinates of the th point, are the coordinates of the th point, is the bandwidth parameter of the Gaussian kernel function.

[0056] Inferring the hierarchical relationship through 3-layer GNN iteration includes: Constructing the first-layer GNN, the second-layer GNN, and the third-layer GNN. The first-layer GNN is used to process the original node features. The second-layer GNN is used to input the node embeddings of the first-layer GNN into the coarsening module. The third-layer GNN aggregates the clusters of the second layer to form semantic units, and maps the clusters of the target layer back to the node relationships of the original graph through the full expansion function.

[0057] Among them, the over-fully extended function is the core technology in the hierarchical GNN interpretation method (such as STFExplainer) that maps the clusters of the target layer back to the node relationships of the original graph. Its core goal is to associate high-level graph structures (such as clustering clusters) with the underlying node relationships of the original graph to preserve cross-level semantic consistency. In a hierarchical graph neural network (HierarchicalGNN), graph data is usually decomposed into multi-level clustering structures (for example, represented by a clustering tree). High-level (such as cluster-level) interpretations need to be associated with the local structure of the original graph (node-level), but traditional methods cannot directly handle this cross-level mapping, resulting in a lack of traceability and interpretability for high-level interpretations. The fully extended function associates the abstract high-level cluster interpretations with the specific structure of the original graph through a mathematical mapping, solving the cross-level semantic discontinuity problem in hierarchical GNN interpretations.

[0058] In an exemplary embodiment of the present invention, the processing sent to the visual channel and the text channel includes: Determine whether the recognition is successful. If the first recognition result and the second recognition result are successfully obtained, perform feature fusion; If the recognition is not successful, adjust the detection threshold, use SIFT feature matching, and determine again whether the recognition is successful. If the recognition is successful, save the current detection threshold. If the recognition is not successful, send an alarm signal and output an error report.

[0059] Among them, SIFT feature matching is an image matching algorithm based on local features, and its goal is to achieve stable detection and matching of key points in images under different scales, rotations, illuminations, and partial occlusions.

[0060] Construct a Gaussian pyramid: Convolve the original image with a Gaussian kernel with different scale parameters σ to generate a multi-scale image sequence, simulating the observation characteristics of the human eye for near and far objects. Fit a three-dimensional quadratic function through Taylor expansion to refine the position and scale of the extreme points and improve the positioning accuracy. Centered on the key points, statistically calculate the gradient magnitude and direction of the pixels in the neighborhood (usually divided into 8 directions). The main direction is the peak of the histogram, and the secondary directions are other directions above 80% of the peak. Divide the 16×16 neighborhood around the key points into 4×4 sub-regions, calculate the 8-direction gradient histogram for each sub-region to form a 128-dimensional vector, and perform normalization processing: perform L2 normalization on the descriptor and truncate it (such as limiting the maximum value to 0.2) to reduce the interference of illumination changes.

[0061] In an exemplary embodiment of the present invention, the script generator respectively generates executable test scripts including: The operation inference finite state machine includes determining whether the type of the currently input element is a login button or an input box. If it is a login button, it checks the adjacent text and generates a click operation script. If it is an input box, it searches for the label associated with the input box and generates a type operation script; Write an assertion generation algorithm, where the assertions include text assertions, structural assertions, and visual assertions.

[0062] Among them, the Click operation script is mainly used to simulate or respond to the user's click behavior, and the Type operation script is used for command-line parameter type processing or system command type judgment.

[0063] This overall solution can be divided into: Input layer: Supports multi-resolution screenshot / video stream input; Processing layer: Four-stage processing pipeline (preprocessing, recognition, parsing, generation); Output layer: Cross-frame executable script and test report.

[0064] And an example is given for the above content: Taking cross-resolution login test generation as the background: Input: Screenshot of the login page with a resolution of 1920x1080 (including username, password input boxes, and login button), test requirement: Verify successful login and redirection.

[0065] Processing process: Preprocessing, scaling to the reference resolution, and eliminating interface noise; Element recognition: Locate 2 input boxes (type: text), 1 button (type: button); OCR extraction: Label texts of "username" and "password"; Layout analysis: Construct a Z-shaped visual flow from the input box to the button.

[0066] Output script.

[0067] Taking responsive layout compatibility test as the background: Input: Comparison screenshots of the same page on the mobile side (375x812) and the desktop side (1440x900) Processing: Element relationship mapping, establishing a cross-resolution element correspondence table, and generating an adaptive positioning strategy.

[0068] In addition, in each embodiment of the present invention, each functional unit can be integrated in one processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0069] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.

[0070] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A cross-platform GUI test script generation method based on multimodal visual recognition, characterized in that: include: Acquire basic data to be processed, the basic data including image data or video data, and pre-process the basic data; The preprocessed basic data are sent to the visual channel and the text channel for processing respectively. The visual channel performs element detection on the basic data through the YOLOv5 model to obtain a first recognition result, and the text channel performs text recognition on the basic data through the CRNN-OCR model to obtain a second recognition result; The first recognition result and the second recognition result are feature fused, the layout relationship is parsed through CNN, the obtained parsed results are output to the script generator, and executable test scripts are generated respectively through the script generator.

2. The cross-platform GUI test script generation method based on multimodal visual recognition according to claim 1 is characterized in that: The preprocessing of basic data includes: The basic data is resolution-standardized and then output as first basic data, the first basic data is noise-eliminated and then dynamic element filtering is performed to output as second basic data; Perform image enhancement on the second basic data and output the processed basic data.

3. The cross-platform GUI test script generation method based on multimodal visual recognition according to claim 2 is characterized in that: The dynamic element filtering includes: Setting static conditions and dynamic conditions, wherein the static conditions include setting filtering thresholds based on numerical ranges or string matching, and the dynamic conditions include obtaining input or real-time data generation conditions; The data set is traversed and a conditional judgment model is set. The conditional judgment model includes logical judgment on the attributes of the elements to confirm whether the attributes of the elements are within the set filtering threshold. After the judgment is completed, the data set obtained by filtering in accordance with the conditions is output.

4. The cross-platform GUI test script generation method based on multimodal visual recognition according to claim 3 is characterized in that: The feature fusion of the first recognition result and the second recognition result includes cross-modal attention alignment, including: Perform bilinear attention interaction: Perform hierarchical attention aggregation, including using a multi-head mechanism to parallelize multi-granularity associations: Perform residual feature fusion: Set the loss function and perform contrast loss optimization; In the formula, is the cross-modal similarity score between the first recognition result and the second recognition result, a device for the first recognition result, is the first words or embedding vectors, is the multi-head attention function, is the vector concatenation function, is the calculation result of the 1st to kth attention heads, is the linear transformation matrix of the multi-head attention output, For the first recognition result after residual feature fusion, is the initial first recognition result, is the layer normalization function, Each row corresponds to the association strength of a visual area with all text words, is the second recognition result, and are projection matrices, To obtain the second recognition result after residual feature fusion, for The transpose of .

5. The cross-platform GUI test script generation method based on multimodal visual recognition according to claim 4 is characterized in that: The feature fusion comprises: Performing vector normalization on the first recognition result and the second recognition result; Concatenate the first recognition result and the second recognition result into a joint vector: In the formula, is the fused joint vector, To calculate the weight.

6. The cross-platform GUI test script generation method based on multimodal visual recognition according to claim 4 is characterized in that: The layout relationship parsing by CNN includes: A weighted spatial relationship matrix is ​​constructed, and the hierarchical relationship is iteratively inferred through a three-layer GNN to output a DOM tree with semantic annotations. The DOM tree includes a parent node and the first input box child node, the second input box child node, and the login button child node associated with the parent node.

7. The cross-platform GUI test script generation method based on multimodal visual recognition according to claim 6 is characterized in that: The constructing of a weighted spatial relationship matrix comprises: In the formula, is the spatial relationship matrix Line The elements of the column, is the natural exponential function, and For the The coordinates of the points, For the The coordinates of the points, is the bandwidth parameter of the Gaussian kernel function.

8. The cross-platform GUI test script generation method based on multimodal visual recognition according to claim 7 is characterized in that: The inference of hierarchical relationships through three-layer GNN iterations includes: The first layer GNN, the second layer GNN and the third layer GNN are constructed. The first layer GNN is used to process the original node features, the second layer GNN is used to embed the nodes of the first layer GNN into the input coarsening module, and the third layer GNN aggregates the clusters of the second layer to form semantic units, and maps the clusters of the target layer back to the node relationship of the original graph through the full expansion function.

9. The cross-platform GUI test script generation method based on multimodal visual recognition according to claim 8, characterized in that: The sending to the visual channel and the text channel processing includes: Determine whether the recognition is successful. If the first recognition result and the second recognition result are obtained successfully, perform feature fusion; If the recognition is not successful, the detection threshold is adjusted, SIFT feature matching is used, and it is judged again whether the recognition is successful. If the recognition is successful, the current detection threshold is saved. If the recognition is not successful, an alarm signal is sent and an error report is output.

10. The cross-platform GUI test script generation method based on multimodal visual recognition according to claim 9, characterized in that: The script generator generates executable test scripts respectively including: Operation reasoning finite state machine, including judging whether the element type of the current input is a login button or an input box. If it is a login button, check the adjacent text and generate a click operation script. If it is an input box, find the label associated with the input box and generate a type operation script. An assertion generation algorithm is written, wherein the assertions include textual assertions, structural assertions, and visual assertions.

Citation Information

Patent Citations

  • Long document classification method and device based on hierarchical multi-granularity interaction graph convolutional network

    CN116522195A

  • Detection and identification method for GUI (Graphical User Interface) control of mobile equipment and test script generation method

    CN117853712A

  • Cross-platform burying-point-free data acquisition and intelligent circle selection rule generation method and system

    CN119537158A