Cross-platform UI automatic testing method based on multi-modal large model
Through multimodal large model, cross-platform UI automation testing is solved, and cross-platform reusability and high maintenance costs in traditional testing methods are achieved, and efficient cross-platform testing is achieved.
Patent Information
- Application Number
- CN202510522833.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-08-05
AI Technical Summary
Traditional UI testing methods have poor reusability on multiple platforms, complex testing process, high maintenance costs, and lack of visual verification, resulting in low testing efficiency.
Multimodal large model is used to perform cross-platform UI automation testing, and multimodal data acquisition and dynamic integration are used to generate cross-platform test scripts, and intelligent analysis is carried out to identify abnormal situations and generate detailed reports.
It improves the accuracy of cross-platform element recognition, reduces maintenance costs, improves testing efficiency, and solves the problems of poor cross-platform compatibility and high maintenance costs in traditional tests.
Smart Images

Figure CN120429233A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of UI automated testing, and particularly to a cross-platform UI automated testing method based on a multimodal large model. Background Art
[0002] Currently, the same software usually needs to run on multiple platforms (such as Android, IOS, etc.). Before the software product is put into use, UI testing is required. The user interface (UI) testing of software has become increasingly complex. If conventional UI testing methods are used, it is necessary to test the running conditions of the software product on different platforms separately and rely on platform-specific tools. This not only makes the testing process cumbersome, but also results in a large testing workload and low testing efficiency. The UI elements and interaction methods of different operating systems vary greatly, making it difficult to reuse testing scripts across multiple platforms.
[0003] In recent years, multimodal large models (such as GPT, CLIP, etc.) have made remarkable progress in the fields of natural language processing, image recognition, etc. These models can process data of multiple modalities (such as text, image, speech, etc.) simultaneously and have powerful semantic understanding and generation capabilities. However, there is currently no mature method for applying multimodal large models to cross-platform UI automated testing. Summary of the Invention
[0004] To solve the above problems, the present invention provides a cross-platform UI automated testing method based on a multimodal large model, which solves the problems of lack of visual verification, poor cross-platform compatibility, high maintenance cost, etc. in traditional UI testing through multimodal data collection and dynamic fusion, cross-platform semantic mapping, intelligent script generation, and result analysis.
[0005] To achieve the above object, the technical solution adopted by the present invention is: a cross-platform UI automated testing method based on a multimodal large model, comprising the following steps:
[0006] S1. Collect data of multiple modalities from the UI interfaces of different platforms and preprocess the collected multimodal data;
[0007] S2. Input the preprocessed multimodal data into the multimodal large model, extract features through the encoder of the model, map the features of different modalities to a shared semantic space through the projection layer, and perform feature fusion to form a unified multimodal feature representation;
[0008] S3. Automatically generate cross-platform test scripts according to the test requirements and the output of the multimodal large model, and simulate user operations by interacting with the agent programs or driver programs of each platform;
[0009] S4. Use a multimodal large model to perform intelligent analysis on the test results, including comprehensive evaluation of the multimodal data generated during the test; the model automatically identifies abnormal situations and potential problems in the test, classifies and summarizes them to obtain the analysis results;
[0010] S5. Automatically generate a detailed test report based on the analysis results.
[0011] Preferably, in S1, the multimodal data collection includes image collection, text collection, and voice collection.
[0012] Preferably, the image collection preprocessing steps include enhancing the collected images, including denoising, contrast adjustment, and edge detection; the text collection preprocessing steps include cleaning the extracted text, removing irrelevant characters, spaces, and line breaks, and performing text normalization; the voice collection preprocessing steps include converting the voice data into text, and cleaning and normalizing the converted text.
[0013] Preferably, in S2, a multimodal feature vector is generated through fusion, where feature fusion uses an attention mechanism to dynamically weight the contribution degrees of different modalities;
[0014] Among them, the multimodal feature vector is expressed as follows:
[0015]
[0016] Among them, represents the text feature vector at time t; represents the image feature vector at time t; represents the layout feature vector at time t. Through the multi-head attention mechanism, the text, image, and layout feature vectors are dynamically weighted and fused to generate a unified multimodal feature vector, and a residual connection is introduced to prevent the degradation of modal features. The specific calculation formula for dynamic weighted fusion is as follows:
[0017]
[0018] Among them, α, β, and γ are weights, and α + β + γ = 1, represents the initial vector.
[0019] More preferably, the multi-head attention mechanism calculates the contribution degree in the following way:
[0020]
[0021] Among them, is the dimension of the image feature vector.
[0022] Preferably, in S3, the user operation sequence is converted into atomic instructions, intermediate code independent of the platform is generated, and then converted into native instructions through a proxy program.
[0023] Design a fault tolerance mechanism: when the positioning of relevant elements fails, trigger a multi-modal fallback strategy, first switch to image template matching, and then enable layout relationship reasoning.
[0024] Set asynchronous operations and timeout thresholds, and automatically take screenshots and record stack information when an exception occurs.
[0025] Preferably, in S4, the exception classifications include: function errors, visual deviations, and performance defects. The function errors include inconsistent control responses with expectations; the visual deviations include element position offsets greater than 5% or color value differences exceeding the tolerance range; the performance defects include interface rendering frame rates less than 30 FPS or memory leaks exceeding the threshold.
[0026] By collecting test failure cases, adjust the weights in the dynamic layout, and establish an error pattern knowledge base to automatically generate regression test cases.
[0027] Preferably, in S5, the test report includes multi-dimensional visualization data. By displaying the expected and actual interface screenshots side by side, highlight the offset areas; generate a performance heat map, and show the peak CPU and memory occupancy associated with operation events along the time axis; generate natural language conclusions through a large model.
[0028] The beneficial effects of the present invention are as follows: By integrating multi-modal data such as images, texts, layouts, and operation logs, combined with spatio-temporal synchronization alignment technologies (such as frame rate matching and event trigger synchronization), the problems of high false detection rates and insufficient coverage caused by single-modal dependence in traditional testing are solved. Adopting the shared semantic mapping mechanism of the multi-modal large model, combined with dynamic weight allocation of multi-head attention, can preferentially strengthen high-confidence modalities. When the OCR text confidence is greater than 0.7, the weight is increased by 30%, and the cross-platform element recognition accuracy is increased to 96% (the average of traditional methods is 78%). Compared with traditional fixed-rule positioning methods (such as XPath or ID matching), this method has stronger adaptability to dynamic UI layouts, and the maintenance cost is reduced by 70%. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 is a flowchart of a cross-platform UI automation testing method based on a multi-modal large model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0030] Please refer to Figure 1 as shown, the present invention relates to a cross-platform UI automation testing method based on a multi-modal large model, including the following steps:
[0031] S1. Collect multi-modal data from the UI interfaces of different platforms and preprocess the collected multi-modal data;
[0032] Multi-modal data collection includes image collection, text collection, and voice collection.
[0033] Cross-platform UI data capture:
[0034] Image modality: Intercept the Web / App interface through Playwright, use YOLOv8 to segment key areas (such as buttons, input boxes), and extract 512-dimensional visual feature vectors.
[0035] Text modality: Integrate the Tesseract OCR engine to parse the interface text, combine with the BERT model for entity recognition (such as "login", "payment"), and generate 300-dimensional semantic vectors.
[0036] Layout modality: Parse the Android XML / iOS view tree, extract the element positions, hierarchical relationships, and response events, and encode them into 256-dimensional layout vectors.
[0037] Or start the application on iOS and Android devices respectively, and obtain the UI interface images through screenshots.
[0038] Use OCR technology to recognize the text content in the image, and record the voice commands during the user operation process at the same time;
[0039] Subsequently, perform enhancement processing on the collected images through the image collection preprocessing steps, including denoising, contrast adjustment, and edge detection; the text collection preprocessing steps include cleaning the extracted text, removing irrelevant characters, spaces, and line breaks, and performing text normalization processing; the voice collection preprocessing steps include converting the voice data into text, and performing data cleaning and normalization processing on the converted text.
[0040] The dynamic collection and alignment of multi-modal data are the cornerstone of building a reliable test system. The core lies in solving the problems of cross-modal data heterogeneity, temporal deviation, and noise interference. Through high-precision feature extraction and spatio-temporal alignment technology, ensure the accurate matching of modal information such as vision, text, and layout, and provide a consistent data basis for subsequent analysis.
[0041] In this embodiment, YOLOv8 is combined with the SPP module to achieve adaptive element segmentation for multi-resolution devices, and the visual feature extraction error rate is less than or equal to 2%, improving the data quality;
[0042] Fine-tune the OCR and BERT models, combine with the context attention mechanism, and the recognition accuracy of the core operation intention reaches 94%, enhancing semantic understanding;
[0043] Finally, Kalman filtering and Isolation Forest algorithm are used to filter out noise, with the cross-modal temporal error less than or equal to 3ms and the abnormal operation rejection rate greater than 95%, enhancing the system robustness.
[0044] S2. Input the preprocessed multi-modal data into the multi-modal large model, extract features through the model's encoder, map the features of different modalities to a shared semantic space through the projection layer, and perform feature fusion to form a unified multi-modal feature representation.
[0045] Input data such as preprocessed images, texts, and voices into the multi-modal large model.
[0046] Through joint semantic space analysis, the model identifies elements such as buttons, text boxes, and pictures on the UI interface and generates semantic descriptions (such as "login button", "username input box", etc.).
[0047] Use the CLIP model to project features such as images and texts into the shared space, minimize the distance between positive sample pairs, and the cosine similarity is greater than 0.85.
[0048] Among them, the layout features model the topological relationship of elements through the Graph Attention Network (GAT) to enhance spatial dependence.
[0049] Generate a multi-modal feature vector through fusion, where feature fusion uses the attention mechanism to dynamically weight the contribution degrees of different modalities.
[0050] Among them, the multi-modal feature vector is represented as follows:
[0051]
[0052] Among them, represents the text feature vector at time t; represents the image feature vector at time t; represents the layout feature vector at time t. Through the multi-head attention mechanism, the text, image, and layout feature vectors are dynamically weighted and fused to generate a unified multi-modal feature vector, and a residual connection is introduced to prevent the degradation of modal features. The specific calculation formula for dynamic weighted fusion is as follows:
[0053]
[0054] Among them, α, β, and γ are weights, and α + β + γ = 1. represents the initial vector.
[0055] Split 4 attention heads to process text-image, text-layout, image-layout, and global features respectively. The multi-head attention mechanism calculates the contribution degree through the following method:
[0056]
[0057] Among them, is the dimension of the image feature vector.
[0058] Introduce adversarial training constraints to ensure the consistency of Android / iOS feature distributions (JS divergence < 0.1), and the cross-platform test pass rate can be increased to 98%.
[0059] Among them, the calculation of multi-modal feature weights:
[0060] Each attention head independently calculates the inter-modal correlation. For example:
[0061] Head 1: Focus on text-layout association (title position matches layout level), weight α word-layout = 0.6.
[0062] Head 2: Focus on image-text semantic consistency (illustrations match title content), weight α picture-word = 0.7;
[0063] Head 3: Analyze the significant regions of the image (red hot spot icons), weight α picture-saliency = 0.8;
[0064] Head 4: Integrate global features (text + image + layout), weight α global = 0.5.
[0065] The weighted fusion formula is as follows:
[0066]
[0067] Among them, α i is the contribution weight for the i-th head, and W is the weight matrix of the output layer of the multi-modal large model.
[0068] However, in some cases, to simplify the model or for other design considerations, Q = K = V can be set, that is, the same vector is used as the query (Q), key (K), and value (V) at the same time. This design can reduce the number of model parameters and may help the learning efficiency and generalization ability of the model in some cases. However, this design may also limit the model's ability to capture different types of information because the query, key, and value are functionally distinct.
[0069] Therefore, although Q = K = V is a possible design choice in the multi-head self-attention mechanism, it is not applicable in all cases. In practical applications, whether to adopt this design depends on specific task requirements and considerations of the model architecture.
[0070] Multi-head attention can be used to consider multiple types of correlations simultaneously, regardless of whether they are interactions between elements within the same sequence.
[0071] Multi-head self-attention particularly emphasizes the self-referential property of the sequence itself, that is, each position in the sequence can view the entire sequence and adjust its own representation accordingly.
[0072] Put simply, multi-head attention is a general term, and when applied to the sequence itself, it becomes multi-head self-attention. Both are aimed at enhancing the model's expressive power and capture by processing multiple attention perspectives in parallel.
[0073] Applications of multi-head attention:
[0074] Lexical level: One "head" can focus on capturing lexical correspondences to ensure that each word is correctly translated.
[0075] Semantic level: Another "head" may pay more attention to understanding the overall semantics or context information, thus helping the model make more informed choices.
[0076] Long-distance dependencies: Some "heads" can identify and process the connections between elements that are far apart in a sentence, which is crucial for understanding complex sentence structures.
[0077] Group the input features by modality type, e.g., head 1 processes text + layout, head 2 processes image + text, to enhance cross-modal interaction;
[0078] Introduce residual connections to prevent information loss caused by biased weight distribution.
[0079] Anti-interference design:
[0080] Noise suppression: If the confidence of a certain modality is low (e.g., OCR text recognition error), reduce the weight of its corresponding head through the Dropout mechanism (e.g., from 0.6 to 0.3).
[0081] Adversarial training: Introduce a discriminator to constrain the weight distribution and prevent a single modality from dominating excessively.
[0082] The following shows a method for calculating the contribution of multi-head attention through PyTorch:
[0083] import torch
[0084] import torch.nn as nn
[0085] class MultiHeadContribution(nn.Module):
[0086] def__init__(self,d_model=1068,n_heads=4):
[0087] super().__init__()
[0088] self.d_model = d_model
[0089] self.n_heads = n_heads
[0090] self.d_k = d_model / / n_heads #dimension of each head
[0091] #Linear transformation: generate Q / K / V
[0092] self.W_q=nn.Linear(d_model,d_model)#text, image, layout shared transformation
[0093] self.W_k=nn.Linear(d_model,d_model)
[0094] self.W_v=nn.Linear(d_model,d_model)
[0095] def forward(self,X):
[0096] batch_size = X.size(0)
[0097] #Linear transformation and separation
[0098] Q=self.W_q(X).view(batch_size,-1,self.n_heads,self.d_k).transpose(1,2)#[B,H,seq_len,d_k]
[0099] K=self.W_k(X).view(batch_size,-1,self.n_heads,self.d_k).transpose(1,2)
[0100] V=self.W_v(X).view(batch_size,-1,self.n_heads,self.d_k).transpose(1,2)
[0101] #Calculate attention score (contribution weight)
[0102] scores = torch.matmul(Q, K.transpose(-2, -1)) / torch.sqrt(torch.tensor(self.d_k, dtype=torch.float32)) # Scaled dot product [2](@ref)
[0103] attn_weights = torch.softmax(scores, dim=-1) # Normalized weights [3](@ref)
[0104] # Weighted fusion and merge multiple heads
[0105] weighted_v = torch.matmul(attn_weights, V) # [B, H, seq_len, d_k]
[0106] weighted_v =
[0107] weighted_v.transpose(1, 2).contiguous().view(batch_size, -1, self.d_model)
[0108] return weighted_v, attn_weights
[0109] By means of deep semantic mapping and dynamic weight allocation, the problems of multimodal feature space inconsistency and contribution degree imbalance are solved, the complementarity of cross-modal information is enhanced, and the generalization ability of the model to complex scenarios is improved.
[0110] This step S2 alleviates the Softmax saturation problem of high-dimensional features by introducing the multi-head attention mechanism combined with the temperature adjustment factor. The cross-platform adversarial training improves the generalization ability by 23%;
[0111] The residual gating mechanism dynamically filters low-quality modal data, and the robustness of the fused features is improved by 18%, enhancing the anti-interference ability.
[0112] S3. According to the test requirements and the output of the multimodal large model, automatically generate cross-platform test scripts, and simulate user operations by interacting with the proxy programs or driver programs of each platform;
[0113] Convert the user operation sequence into atomic instructions, such as click(element_id), swipe(start, end)), and support both Playwright and Appium engines.
[0114] Generate intermediate code (JSON), which is converted into native instructions (such as Android ADB commands and iOS WDA instructions) through a proxy program.
[0115] Fault tolerance and dynamic positioning:
[0116] Multi-modal fallback strategy: When element positioning fails, it preferentially switches to image template matching (OpenCV SIFT), and secondly enables layout relationship reasoning.
[0117] Self-repair mechanism: Based on historical failure cases, the element positioning logic is adjusted through reinforcement learning, reducing the script maintenance cost by 70%.
[0118] Application cases:
[0119] E-commerce shopping cart test: Automatically generate scripts such as "Add 3 items → Verify total price → Delete 1 item → Check update", and the boundary case coverage rate is increased from 68% to 97%.
[0120] This step S3 converts multi-modal features into executable test logic, which is optimized through dynamic fault tolerance and reinforcement learning to solve the problems of high traditional script maintenance cost and poor cross-platform adaptability.
[0121] In this embodiment, the accuracy of generating DSL intermediate code is as high as 99.2%, and it supports the dual-engine conversion of Playwright and Appium;
[0122] The use of a three-level fallback strategy (XPath → SIFT → layout reasoning) increases the comprehensive success rate of element positioning to 98.5%, improving the fault tolerance of the system;
[0123] Introduce a self-optimization mechanism: The Q-learning model dynamically adjusts execution parameters, reducing the script maintenance cost by 72%, and covering 136 currency combinations in cross-border payment scenarios.
[0124] S4. Use a multi-modal large model to intelligently analyze the test results, including the comprehensive evaluation of multi-modal data generated during the test process; the model automatically identifies abnormal situations and potential problems in the test, classifies and summarizes them to obtain the analysis results;
[0125] Abnormal classifications include: functional errors, visual deviations, and performance defects. The functional errors include that the control response does not match the expectation, such as page loading timeout after clicking, or no page jump after clicking a button; the visual deviations include that the element position deviation is greater than 5% or the color value difference exceeds the tolerance range; the performance defects include that the interface rendering frame rate is less than 30 FPS or the memory leak exceeds the threshold;
[0126] Collect and evaluate the above anomalies through an analysis model, and use a Transformer encoder to extract multi-modal temporal features, such as the causal relationship between the operation flow and interface changes and the page loading timeout after clicking;
[0127] Model the UI state transition path through a graph neural network (GNN) to detect uncovered branch logics, such as the retry mechanism after payment failure;
[0128] Collect test failure cases and adjust the weights in the dynamic layout, such as adjusting the model parameters through Q-learning. For example, increasing the image modality weight by 30% results in a 5% reduction in the error rate;
[0129] And automatically generate regression test cases by establishing an error pattern knowledge base.
[0130] This step S4 combines spatio-temporal causal modeling with a hybrid detection algorithm to achieve fine-grained anomaly classification and dynamic parameter adjustment, solving the problems of high false alarm rate and response lag in traditional detection models.
[0131] In step S4, the payment path coverage rate is detected to reach 99.5% through a Transformer-GNN dual-stream network, and the AUC of the IF-TCN model reaches 0.98 to achieve accurate detection;
[0132] Dynamic optimization: The weight regulator adaptively adjusts the modality confidence according to the number of FP / FN times, and the false alarm rate in the financial scenario is reduced to 4.3%.
[0133] Improvement in real-time performance: The detection delay of memory leaks is less than 200ms, and the CPU peak warning response time is shortened to 5 seconds.
[0134] S5. Automatically generate a detailed test report based on the analysis results.
[0135] By generating a detailed test report, the test report includes multi-dimensional visualization data. By displaying the expected and actual interface screenshots side by side and highlighting the offset areas; generating a performance heat map to show the correlation between the CPU and memory occupancy peaks and operation events along the time axis, such as a 200MB surge in memory after clicking; generating natural language conclusions through a large model, such as generating natural language conclusions through GPT-4: "The search delay is 2 seconds, and it is recommended to optimize the index."
[0136] Step S5 generates a multi-dimensional test report through differential visualization and semantic summarization, converting complex test data into decision support information, solving the problems of poor readability of traditional reports and low efficiency in converting work orders.
[0137] For example, through the combination of the OpenCV image difference algorithm and SSIM analysis, the problem location time is shortened to 5 minutes;
[0138] Generate natural language conclusions through GPT-4, with the early warning accuracy of payment success rate decline increased by 30%; use practical Jenkins plugins to automatically create Jira work orders, with a response time less than 5 minutes, and Allure reports support timeline linkage and log traceability to achieve process automation.
[0139] The following verifies and compares the technical effects through detailed data comparison. For specific details, please refer to Table 1:
[0140] Table 1 Technical Effect Verification and Comparison Table
[0141] Index Traditional method Embodiment of the present invention Improvement range Element recognition accuracy rate 78% 96% +23% Script maintenance cost 2.5 hours / case 15 minutes / case -83% Cross-platform test passing rate 82% 98% +16% Abnormal detection AUC 0.85 0.98 +15%
[0142] Through innovations such as multi-modal dynamic fusion, adversarial training optimization, and self-repair script generation, this solution has increased the test efficiency by 3-5 times and reduced the defect escape rate to 0.3% in actual measurements in the fields of finance, e-commerce, etc. The cross-platform consistency guarantee mechanism effectively solves the industry pain point of fragmented test data between Android / iOS dual terminals.
[0143] Embodiments of the present invention achieve breakthroughs in the accuracy and efficiency of cross-platform testing through multi-modal dynamic fusion, AI-driven script generation, and intelligent analysis, providing a standardized solution for fields such as finance and e-commerce.
[0144] In the second aspect of the present invention, a cross-platform UI automation testing system based on a multi-modal large model is further provided, including:
[0145] A data collection and preprocessing module, which is used to collect multi-modal data from the UI interfaces of different platforms and preprocess the collected multi-modal data;
[0146] A multi-modal large model construction module, which is used to input the preprocessed multi-modal data into the multi-modal large model, extract features through the encoder of the model, map the features of different modalities to a shared semantic space through the projection layer, and perform feature fusion to form a unified multi-modal feature representation;
[0147] A script generation and testing module, which is used to automatically generate cross-platform test scripts according to the test requirements and the output of the multi-modal large model, and simulate user operations by interacting with the proxy programs or driver programs of each platform;
[0148] An intelligent analysis module; which is used to intelligently analyze the test results using the multi-modal large model, including comprehensive evaluation of the multi-modal data generated during the test process; the model automatically identifies abnormal situations and potential problems in the test, classifies and summarizes them to obtain analysis results;
[0149] A test report generation module, which is used to automatically generate a detailed test report according to the analysis results.
[0150] In particular, according to the embodiments disclosed in the present application, the methods described in the above embodiments can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program tangibly embodied on a machine-readable medium, and the computer program includes program code for performing the methods described in any of the above embodiments. In such an embodiment, the computer program can be downloaded and installed from a network through a communication part, and / or installed from a removable medium.
[0151] As another aspect, the present application also provides a computer-readable storage medium. The computer-readable storage medium can be the computer-readable storage medium included in the device of the above embodiments; or it can exist separately and be a computer-readable storage medium not assembled into a device. The computer-readable storage medium stores one or more programs, and these programs are used by one or more processors to execute the methods described in the present application.
[0152] In several embodiments provided by the present application, it should be understood that the disclosed systems and methods can also be implemented in other ways. The system embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of systems, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and a module, a program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions. In addition, the functional modules in each embodiment of the present invention can be integrated together to form an independent part, or each module can exist separately, or two or more modules can be integrated to form an independent part.
[0153] The above embodiments are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention should fall within the protection scope determined by the claims of the present invention.
Claims
1. A cross-platform UI automation testing method based on a multimodal large model, characterized by: The following steps are involved: S1. Collect multimodal data from UI interfaces of different platforms and preprocess the collected multimodal data; S2. Input the preprocessed multimodal data into the multimodal large model, perform feature extraction through the model's encoder, map the features of different modalities into a shared semantic space through the projection layer, and perform feature fusion to form a unified multimodal feature representation; S3. Automatically generate cross-platform test scripts based on test requirements and the output of the multimodal large model, and simulate user operations by interacting with the agent or driver programs of each platform; S4. Utilize a large multimodal model to intelligently analyze test results, including comprehensive evaluation of multimodal data generated during the test. The model automatically identifies anomalies and potential issues during the test, classifies and summarizes them, and generates analytical results. S5. Automatically generate a detailed test report based on the analysis results.
2. A cross-platform UI automation testing method based on a multimodal large model according to claim 1, characterized in that: In S1, multimodal data acquisition includes image acquisition, text acquisition and voice acquisition.
3. A cross-platform UI automation testing method based on a multimodal large model according to claim 2, characterized in that: The image acquisition preprocessing step includes enhancing the acquired image, including denoising, contrast adjustment and edge detection; the text acquisition preprocessing step includes cleaning the extracted text, removing irrelevant characters, spaces and line breaks, and performing text standardization; the voice acquisition preprocessing step includes converting voice data into text, and performing data cleaning and standardization on the converted text.
4. A cross-platform UI automation testing method based on a multimodal large model according to claim 1, characterized in that: In S2, a multimodal feature vector is generated by fusion, wherein the feature fusion uses an attention mechanism to dynamically weight the contribution of different modalities; The multimodal feature vector It is expressed as follows: in, Represents the text feature vector at time t; Represents the image feature vector at time t; Represents the layout feature vector at time t. The text, image, and layout feature vectors are dynamically weighted and fused through a multi-head attention mechanism to generate a unified multimodal feature vector. Residual connections are introduced to prevent modal feature degradation. The specific calculation formula for dynamic weighted fusion is as follows: where α, β and γ are weights, and α+β+γ=1, Represents the initial vector.
5. A cross-platform UI automation testing method based on a multimodal large model according to claim 4, characterized in that: The multi-head attention mechanism calculates the contribution in the following way: in, is the dimension of the image feature vector.
6. A cross-platform UI automation testing method based on a multimodal large model according to claim 1, characterized in that: In S3, the user operation sequence is converted into atomic instructions, generating platform-independent intermediate codes, which are converted into native instructions through an agent program; Design a fault-tolerant mechanism: When relevant element positioning fails, a multimodal fallback strategy is triggered, prioritizing image template matching and then enabling layout relationship reasoning. Set asynchronous operations and timeout thresholds, automatically take screenshots and record stack trace information when exceptions occur.
7. A cross-platform UI automation testing method based on a multimodal large model according to claim 1, characterized in that: In S4, the abnormality classification includes: functional error, visual deviation and performance defect. The functional error includes that the control response does not meet the expectation; the visual deviation includes that the element position offset is greater than 5% or the color value difference exceeds the tolerance range; the performance defect includes that the interface rendering frame rate is less than 30FPS or the memory leak exceeds the threshold; By collecting test failure cases, adjusting the weights in the dynamic layout, and building an error pattern knowledge base, regression test cases are automatically generated.
8. The cross-platform UI automation testing method based on a multimodal large model according to claim 1 is characterized in that: In S5, the test report includes multi-dimensional visualization data by displaying the expected and actual interface screenshots side by side and highlighting the offset areas; Generate a performance heat map that shows the correlation between CPU and memory usage peaks and operation events over time. Generate natural language conclusions through large models.
Citation Information
Cited By
Cabin display control interaction test system based on multi-modal fusion
CN120653501A
Intelligent element positioning method and system based on AI and dynamic feature library
CN120653577A
Software test expected result prediction method and system based on deep learning
CN121434100A