Data path consistency detection method for terminal equipment

By extracting the data paths in the terminal device privacy policy and performing consistency detection, the problem of inconsistency between the terminal device application data paths and the privacy policy statements is solved, efficient and automated data path consistency detection is achieved, and user privacy rights are protected.

CN120180490APending Publication Date: 2025-06-20XI AN JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510249825.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The data paths in terminal device applications may be inconsistent with the contents of the Privacy Policy Statement, resulting in the threat of user privacy rights, and the existing technology has failed to effectively solve this problem.

Method used

A data path consistency detection method for terminal devices is proposed. Data paths are extracted through regular expressions and large language models, dynamic testing is optimized using prefix trees, multimodal large language models are automated and dynamic testing, and consistency results are tested through word embedding models.

Benefits of technology

It realizes consistent detection of terminal equipment data paths, improves detection efficiency and accuracy, saves manual review time, and protects users' personal data privacy rights.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180490A_ABST
    Figure CN120180490A_ABST
Patent Text Reader

Abstract

The invention discloses a data path consistency detection method for terminal equipment. The method comprises the following four steps of: performing data path extraction, preprocessing, dynamic testing and consistency result judgment on the terminal equipment: firstly, extracting a data path in a privacy policy text of the terminal equipment by using a regular expression and a large language model; then, establishing a prefix tree for a plurality of data paths to optimize a task sequence of the dynamic test; then, the multi-mode large language model is used for carrying out automatic dynamic testing on the application program of the terminal equipment, and an actual data path is obtained; and finally, comparing the difference between the actual data path and the declared data path through a word embedding model, and obtaining a final consistency check result. Through the method, the data path consistency detection of the terminal equipment can be efficiently and automatically completed, the manpower and time for manually checking the problem are saved, and the personal data privacy right of the user is also guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the problem of data privacy of terminal equipment, and in particular to a dynamic testing framework for data path consistency verification of terminal equipment. Background Art

[0002] While humans enjoy applications on terminal devices, they also rely on them. Mobile terminal applications have brought great convenience to people's lives. The number and types of applications are growing rapidly to meet people's growing needs. Due to a series of privacy leaks and violations in terminal applications, users are increasingly concerned about privacy rights and protection, which has prompted governments to implement stricter privacy protection measures and privacy disclosure regulations. Developers of terminal device applications usually provide personal permission paths in privacy policies to help users exercise their personal rights. However, in the process of users exercising and managing their personal rights, the data path in the privacy policy may be inconsistent with the actual behavior in the terminal application. This is because in actual application development and operation and maintenance, developers are usually not the writers of the privacy policy, or the privacy policy is not updated in a timely manner with the update of the application code.

[0003] Existing research mainly focuses on detecting applications that violate the content of the privacy policy statement to ensure that the personal data collection behavior in the privacy policy is consistent with the actual behavior of the terminal application. However, current work does not pay attention to the consistency between the data path in the privacy policy and the actual behavior of the terminal application. Therefore, the content of this invention will fill this gap. Summary of the invention

[0004] The content of the present invention is to propose a data path consistency detection method for terminal devices to solve the above technical problems. This framework is divided into four steps, which respectively perform data path extraction, preprocessing, dynamic testing and consistency result judgment on the terminal device: first, use regular expressions and large language models to extract the data path in the privacy policy text of the terminal device; then, establish a prefix tree for several data paths to optimize the task sequence of dynamic testing; then, use a multimodal large language model to carry out automated dynamic testing of the terminal device application to obtain the actual data path; finally, use a word embedding model to compare the actual data path with the declared data path, and obtain the final consistency test result. Through the above method, the data path consistency detection of the terminal device can be completed efficiently and automatically, saving manpower and time for manual review of the problem, and also protecting the user's personal data privacy rights.

[0005] In order to achieve the above object, the present invention adopts the following technical solutions:

[0006] 1. Data path extraction based on regular expressions and large language models, including the following steps:

[0007] 1) Data path extraction based on regular expressions and large language models

[0008] Step S101: Input the privacy policy P of the terminal device, and use regular expressions to extract the statements containing special symbols "[ ], "", -, >" from the privacy policy;

[0009] Step S102: According to the statements obtained in Step S101, and using the prompt L1, input the statement and the prompt together into the large language model, and the large language model outputs the statement containing the data path, where the statement contains the data path and the functional description of the corresponding scenario.

[0010] 2) Test path sequence based on prefix tree

[0011] Step S103: According to the data path obtained in S102, use each passing point in the data path as a node to construct a Trie prefix tree, and at the same time mark the corresponding node of the end point of each data path;

[0012] Step S104: According to the prefix tree obtained in S103, perform a breadth-first traversal of the tree. When the marked node is traversed, output the data path sequence until all nodes on the prefix tree are traversed. Finally, use the sequentially output data path sequences as the test path sequences.

[0013] 3) Dynamic testing of terminal devices based on multi-modal large language models

[0014] Step S105: According to the test path sequence obtained in S104, input the application program of the terminal device to be tested, and use the simulation operation driver to open the application program;

[0015] Step S106: Take a screenshot of the current application program page and obtain the UI element hierarchy file, and perform preprocessing to obtain the preprocessed screenshot and the filtered UI element list;

[0016] Step S107: According to the filtered UI element list obtained by preprocessing in Step S106, use its hierarchy file to generate text information for the UI elements, and use the prompt to merge the preprocessed screenshot and the UI element text information obtained in Step S106 and input them to the multi-modal large language model, and the multi-modal large language model outputs the next application program operation;

[0017] Step S108: Convert the next application operation input obtained in S107 to the driver of the simulation operation, take a screenshot of the application page after the operation, and preprocess the UI element hierarchy file. Use the prompt to merge the input and feed it to the multi-modal large language model to determine whether the operation is successful, and determine whether the application has reached the end of the data path. If so, output the sequence of UI elements operated in this process, use the multi-modal large language model to output the scene description of the final page, and transfer to S109; otherwise, transfer to step S106;

[0018] 4) Verification of data path consistency results based on the word embedding model:

[0019] Step S109: Generate a text description for each UI element according to the UI element sequence obtained in step S107;

[0020] Step S110: Compare the text description of the UI element generated in step S108 and the scene description of the final page with the functional description of the corresponding data path and scene extracted in step S102 to obtain the data path consistency comparison result.

[0021] Further, it is characterized in that in step S102, the large language model is used to extract the data path, specifically:

[0022] Step S201: According to the statements obtained in step S101, use the prompt L1 to input into the large language model to determine whether each statement is a click path. If so, output its corresponding click path and the corresponding scene description, and transfer to step S202; otherwise, do nothing;

[0023] Step S202: Perform unified formatting processing on the click paths obtained in step S201, and the output format is [Path 1]->[Path 2]->…->[Path n].

[0024] Further, it is characterized in that in step S106, the UI element hierarchy file and the screenshot of the application page are preprocessed, specifically:

[0025] Step S301: Traverse the UI element hierarchy file obtained in step S106, extract the UI elements whose clickable and focusable attribute values are true, and select the UI elements whose size does not exceed the threshold. Extract the boundary coordinate information of these UI elements, and finally determine the numbering order of these UI elements according to the traversal order.

[0026] Step S302: Based on the screenshot of the current application page obtained in Step S106, and the UI elements extracted in Step S301, their numbering order, and boundary coordinate information, mark these controls with bounding boxes at the corresponding positions of each UI element on the screenshot, and mark its number directly above the bounding box.

[0027] Further, in Step S107, text information is generated for UI elements using the UI element hierarchy file, specifically:

[0028] Step S401: Filter the text, content-desc, and resource-id attributes from the UI element hierarchy file. If all three attributes are empty, proceed to Step S402; otherwise, concatenate the text information of the three in order as the generated UI element text information.

[0029] Step S402: Obtain the image of the corresponding UI element from the screenshot according to the coordinates of the corresponding UI element obtained in Step S301, and input it into the multimodal large language model to determine whether the UI element is an image of the displayed content or a UI element for switching functions. If it is the former, discard the UI element; otherwise, proceed to Step S403.

[0030] Step S403: Based on the image of the UI element obtained in Step S402, combined with the prompt words, give a functional description of common application UI elements, and input it into the multimodal large language model to let it output the text description corresponding to the UI element. Further, in Step S108, the multimodal large language model outputs a driver program that transforms the next application operation into a simulated operation, specifically:

[0031] Step S501: The multimodal large language model structurally outputs four types of operations, including click, input text, long click, and swipe, with the parameters used being the UI element number, input content, UI element number, and UI element number plus direction and distance respectively.

[0032] Step S502: Based on the operations output by the large language model obtained in Step S501, when the output is a click operation, call the click interface in the simulated operation driver program; when the output is an input text operation, first call the click interface in the simulated operation driver program, and then call the text input interface; when the output is a long click and swipe operation, call the long click and swipe interfaces in the simulated operation driver program.

[0033] Further, it is characterized in that in Step S108, a determination is made on whether a certain operation is successful and whether the application has reached the final path, specifically:

[0034] Step S601: For whether a certain operation is successful, input the screenshots of the application before and after the operation into the multimodal large language model. The large language model outputs three types of results, namely operation successful, operation invalid, and return, based on the differences in the screenshots. Operation successful means that a certain operation achieves the expected effect. Operation invalid means that there are no obvious changes in the screenshots before and after the operation. Return means that the page entered by the current operation cannot reach the target page and needs to return to the previous page;

[0035] Step S602: To determine whether the application has reached the final path, it is necessary to input the screenshot of the current application into the multimodal large language model and combine with prompt words to help the multimodal large language model output the determination result.

[0036] Further, it is characterized in that step S110 judges the data path consistency result, specifically:

[0037] Step S701: Input the text descriptions of the UI element sequences obtained in S109 and the data path texts obtained in step S102 into the word embedding model respectively, output two vectors, calculate the cross-entropy loss value of the vectors. If it is greater than the threshold, output that the data click path is inconsistent; if it is less than the threshold, output that the data click path is consistent;

[0038] Step S702: Input the function descriptions of the final scene obtained in S108 and the function descriptions of the corresponding scenes extracted in step S102 into the word embedding model respectively, output two vectors, calculate the cross-entropy loss value of the vectors. If it is greater than the threshold, the data path scene function descriptions are inconsistent; if it is less than the threshold, output that the data path scene function descriptions are consistent. Further, it is characterized in that the thresholds of step S701 and step S702 are calculated using small samples with labeled tags and extended to large samples without labeled tags, specifically:

[0039] Step S801: Manually label the consistency of a small number of data paths. After performing operations from step 101 to step 109, calculate the average cross-entropy loss value in the data click path samples to obtain the threshold for judging the consistency of the data click path;

[0040] Step S802: Calculate the average cross-entropy loss value in the samples with consistent data path scene functions to obtain the threshold for judging the consistency of the data path scene functions.

[0041] A further improvement of the present invention is that the data path extraction methods of steps S101 and S102 based on regular expressions and large language models utilize the prior law that data paths usually appear accompanied by special symbols, initially filter out irrelevant privacy policy texts, save the extraction cost, and improve the efficiency of the large language model in judging and extracting data paths.

[0042] A further improvement of the present invention lies in that: steps S103 and S104 utilize the traversal of a prefix tree to obtain an optimal dynamic test sequence for the data path.

[0043] A further improvement of the present invention lies in that: steps S105 to S108 utilize a multi-modal large language model to conduct dynamic testing on the application programs of the terminal device. Compared with traditional dynamic testing, which is highly random and inefficient, the dynamic testing based on the multi-modal large language model is highly efficient and can meet the testing requirements by flexibly modifying the prompt words according to specific tasks.

[0044] A further improvement of the present invention lies in that: steps S109 and S110 utilize a word embedding model to measure the consistency between the actual data path and scenario functions of the terminal application program and its privacy policy description. Compared with manual determination of consistency, which requires a large amount of human-labeled samples, the consistency test based on the word embedding model can carry out an evaluation process from small samples to large samples to save the time of labeling samples.

[0045] Compared with the prior art, the present invention has the following advantages:

[0046] 1) The method of the present invention can quickly and automatically detect inconsistent problems in the data path of the terminal device for the data path consistency problem of the terminal device.

[0047] 2) The present invention utilizes regular expressions and a large language model to quickly and accurately extract data paths from the privacy policies of a large number of terminal device application programs.

[0048] 3) The dynamic test module of the present invention utilizes a multi-modal large language model to accurately reach the target scenario without manual intervention and judge the scenario function based on the scenario screenshot.

[0049] 4) The data path consistency evaluation method of the present invention does not require large-scale data annotation. This is because the present invention calculates the threshold from a small sample with labeled tags and applies it to a large sample without labeled tags to quickly judge the data path consistency. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 It is a flowchart of the data path consistency detection method for the terminal device of the present invention;

[0051] Figure 2 It is a flowchart of the data path extraction method of the present invention;

[0052] Figure 3 It is a flowchart of the preprocessing method for the UI element hierarchy file of the present invention;

[0053] Figure 4 This is the flowchart of the method for generating text information for UI elements in the present invention;

[0054] Figure 5 This is the flowchart of the output defined by the large language model in the present invention and the interaction with the terminal device;

[0055] Figure 6 This is the flowchart of the method for judging the running result of the steps of the large language model in the present invention;

[0056] Figure 7 This is the flowchart of the method for discriminating whether the data paths of the terminal device are consistent in the present invention; Detailed implementation manners

[0057] The following will describe in detail a data path consistency detection method for terminal devices proposed by the present invention with reference to the accompanying drawings.

[0058] Refer to Figure 1 As shown, the present invention discloses a data path consistency detection method for terminal devices, including the following steps:

[0059] 1) Data path extraction based on regular expressions and large language models

[0060] Step S101: Input the privacy policy of the terminal device, which can usually be obtained in batches through platforms such as the terminal application market. First, use regular expressions to extract statements containing special symbols "[ ], "", -, >" from the privacy policy

[0061] Step S102: According to the statements obtained in step S101, and using prompt engineering, input this statement and the prompt words into the large language model together. The large language model can be GPT-3, GPT-4, Tongyi Qianwen, etc. The large language model outputs the statement that truly contains the data path, and this statement contains the data path and the functional description of the corresponding scenario (that is, what functions can be achieved through this data path).

[0062] 2) Test path sequence based on prefix tree

[0063] Step S103: According to the data path obtained in S102, use each passing point in the data path as a node to construct a Trie word prefix tree. The growth of the Trie word prefix tree is by judging whether a certain node of the prefix tree exists (), if it exists, continue to traverse the Trie tree. If it does not exist, continue to traverse the next child node of the current node. If all child nodes do not exist, create a new node under the current node. When reaching the element (path) to be detected in a certain data path, mark it by setting a flag bit.

[0064] Step S104: Perform a breadth-first traversal of the prefix tree obtained in S103. When a marked node is traversed, output the data path sequence until all nodes on the prefix tree have been traversed. Finally, use the sequentially output data path sequences as the test path sequences. The number of page switches needs to be determined between the test sequences. This number requires finding the nearest common node of the leaf nodes of the two paths. By obtaining the distance from the leaf node of the current test sequence to this common node, it is judged how many pages should be rolled back during the switching of the test sequence. This distance is different from the distance of traditional tree nodes. The distance will only increase when the path from this common node to the leaf node causes a new page to be generated in the application program of the terminal device.

[0065] 3) Dynamic testing of terminal devices based on multimodal large language models

[0066] Step S105: According to the test path sequence and switching method obtained in S104, input the application program of the terminal device to be tested, and use the simulation operation driver. Taking an Android terminal device as an example, this driver usually opens the application program using existing frameworks such as ADB or UiAutoMator.

[0067] Step S106: Take a screenshot of the current application program page and obtain the UI element hierarchy file, and perform preprocessing to obtain the preprocessed screenshot and the filtered UI element list.

[0068] Step S107: According to the filtered UI element list obtained by preprocessing in Step S106, generate text information for the UI elements using its hierarchy file, and use prompt engineering to merge the preprocessed screenshot and UI element text information obtained in Step S106 and input them into the multimodal large language model. This multimodal large language model can be GPT-4o, Qwen2-vl, etc. The multimodal large language model outputs the next application program operation.

[0069] Step S108: Convert the input of the next application program operation obtained in S107 to the simulation operation driver, take a screenshot of the operated application program page, and preprocess the UI element hierarchy file. Use the prompt words to merge and input them into the multimodal large language model to judge whether the operation is successful, and judge whether the application program has reached the end of the data path. If so, output the UI element sequence of the operations during this process, use the multimodal large language model to output the scene description of the final page, and transfer to S109. Otherwise, transfer to Step S106.

[0070] 4) Verification of data path consistency results based on word embedding models:

[0071] Step S109: Generate a text description for each UI element according to the UI element sequence obtained in Step S107.

[0072] Step S110: Compare the text description of the UI elements generated in step S108 with the scene description of the final page, and compare it with the function description of the corresponding data path and scene extracted in step S102 to obtain the comparison result of data path consistency

[0073] Refer to Figure 2 as shown below:

[0074] Step S201: According to the statements obtained in step S101, input the prompt word L1 into the large language model to determine whether each statement is a data path. If so, output its corresponding click path and corresponding scene description, and then go to step S202; otherwise, do nothing

[0075] Step S202: Perform unified formatting processing on the click paths obtained in step S201, and the output format is [Path 1]->[Path 2]->…->[Path n].

[0076] Refer to Figure 3 as shown below:

[0077] Step S301: Traverse the UI element hierarchy file obtained in step S106, extract the UI elements whose clickable and focusable attribute values are true, select the UI elements whose size does not exceed the threshold, extract the boundary coordinate information of these UI elements, and finally determine the numbering order of these UI elements according to the traversal order.

[0078] Step S302: Based on the screenshot of the current application page obtained in step S106, and the UI elements, their numbering order, and boundary coordinate information extracted in step S301, mark these controls with bounding boxes at the corresponding positions of each UI element on the screenshot, and mark their numbers directly above the bounding boxes.

[0079] Refer to Figure 4 as shown below:

[0080] Step S401: Filter the text, content-desc, and resource-id attributes from the UI element hierarchy file. If all three attributes are empty, go to step S402; otherwise, concatenate the text information of the three in order as the generated UI element text information

[0081] Step S402: According to the coordinates of the corresponding UI elements obtained in step S301, obtain the pictures of the corresponding UI elements from the screenshot, input them into the multi-modal large language model, and determine whether the UI element is a picture of the displayed content or a UI element for switching functions. If it is the former, discard the UI element; otherwise, go to step S403

[0082] Step S403: Based on the image of the UI element obtained in Step S402, combined with the prompt words, provide a functional description of common application program UI elements, and input it into the multimodal large language model to let it output the text description corresponding to the UI element.

[0083] Refer to Figure 5 as shown below, specifically:

[0084] Step S501: The multimodal large language model structurally outputs four types of operations, including click, input text, long click, and swipe. The parameters used are the UI element number, the input content, the UI element number, and the UI element number plus the direction and distance

[0085] Step S502: Based on the operations output by the large language model obtained in Step S501, when the click operation is output, call the click interface in the simulation operation driver; when the input text operation is output, first call the click interface in the simulation operation driver, and then call the text input interface; when the long click and swipe operations are output, call the long click and swipe interfaces in the simulation operation driver.

[0086] Refer to Figure 6 as shown below, specifically:

[0087] Step S601: Regarding whether a certain operation is successful, input the screenshots of the application program before and after the operation into the multimodal large language model. The large language model outputs three types of results, namely operation successful, operation invalid, and return, based on the differences in the screenshots. Operation successful means that a certain operation has achieved the expected effect. Operation invalid means that there is no obvious change in the screenshots before and after the operation. Return means that the page entered by the current operation cannot meet the requirement of reaching the target page and needs to return to the previous page

[0088] Step S602: To determine whether the application program has reached the final path, it is necessary to input the current screenshot of the application program into the multimodal large language model and combine the prompt words to help the multimodal large language model output the determination result.

[0089] Refer to Figure 7 as shown below, specifically:

[0090] Step S701: Input the text description of the UI element sequence obtained in S109 and the data path text obtained in Step S102 into the word embedding model respectively, output two vectors, calculate the cross-entropy loss value of the vectors. If it is greater than the threshold, output that the data click path is inconsistent; if it is less than the threshold, output that the data click path is consistent

[0091] Step S702: Input the function descriptions of the final scenario obtained in S108 and the function descriptions of the corresponding scenarios extracted in step S102 into the word embedding model respectively, output two vectors, calculate the cross-entropy loss value of the vectors. If it is greater than the threshold, it means that the function descriptions of the data path scenarios are inconsistent; if it is less than the threshold, it outputs that the function descriptions of the data path scenarios are consistent. Further, it is characterized in that the calculation of the thresholds in step S701 and step S702 uses small samples with labeled tags, and this is extended to large samples without labeled tags. Specifically:

[0092] Step S801: Manually label the consistency of a small number of data paths. After performing the operations from step S101 to step S109, calculate the average cross-entropy loss value of the word embedding vectors of the data click paths of the samples and the labeled samples to obtain a threshold for judging the consistency of the data click paths. If the average cross-entropy loss value of the word embedding vectors of the data click paths on the large sample is less than this threshold, it should be judged that the data paths are consistent; otherwise, it should be judged as inconsistent.

[0094] Step S802: Calculate the average loss value of the cross-entropy in the samples with consistent scenario functions of the data paths to obtain a threshold for judging the consistency of the scenario functions of the data paths. If the average cross-entropy loss value of the word embedding vectors of the scenario functions of the data paths on the large sample is less than this threshold, it should be judged that the data paths are consistent; otherwise, it should be judged as inconsistent.

Claims

1. A data path consistency detection method for a terminal device, characterized in that: The steps include: 1) Data path extraction based on regular expressions and large language models Step S101: input the privacy policy P of the terminal device, and use regular expressions to extract sentences containing special symbols "[],"", -, >" from the privacy policy; Step S102: According to the sentence obtained in step S101, the sentence and the prompt word L1 are input into the large language model, and the large language model outputs a sentence containing the data path, wherein the sentence contains the data path and the functional description of the corresponding scenario; 2) Test path sequence based on prefix tree Step S103: Based on the data path obtained in S102, each passing point in the data path is used as a node to construct a Trie prefix tree, and the corresponding node of each data path end point is marked; Step S104: according to the prefix tree obtained in S103, perform breadth-first traversal on the tree, and when traversing to the marked node, output the data path sequence until the traversal of all nodes on the prefix tree is completed, and finally use the data path sequence output in sequence as the test path sequence; 3) Dynamic testing of terminal devices based on multimodal large language models Step S105: according to the test path sequence obtained in S104, input the application of the terminal device to be tested, and open the application using the driver program of the simulated operation; Step S106: taking a screenshot of the current application page and obtaining a UI element hierarchy file, and performing preprocessing to obtain a preprocessed screenshot and a filtered UI element list; Step S107: according to the filtered UI element list obtained by preprocessing in step S106, generate text information for the UI element using its hierarchical file, and use prompt words to merge the preprocessed screenshot obtained in step S106 and the UI element text information into the multimodal large language model, and the multimodal large language model outputs the next application operation; Step S108: according to the input of the next application operation obtained in S107, the input is converted to the driver of the simulation operation, a screenshot of the application page after the operation is taken, and the UI element level file is preprocessed, and the prompt word is merged and input to the multimodal large language model to determine whether the operation is successful, and whether the application has reached the end of the data path. If so, the UI element sequence of the operation in the process is output, and the scene description of the final page is output using the multimodal large language model, and the process goes to S109, otherwise it goes to step S106; 4) Data path consistency test based on word embedding model: Step S109: Generate a text description for each UI element according to the UI element sequence obtained in step S107; Step S110: Compare the text description of the UI element generated in step S108 with the scene description of the final page with the functional description of the corresponding data path and scene extracted in step S102 to obtain a data path consistency comparison result.

2. The method according to claim 1, characterized in that The step S102 is specifically as follows: Step S201: Based on the sentences obtained in step S101, the prompt word L1 is input into the large language model to determine whether each sentence is a click path. If so, the corresponding click path and corresponding scene description are output, and the process goes to step S202. Otherwise, no processing is performed. Step S202: Perform unified formatting processing on the click path obtained in step S201, and the output format is [path 1] -> [path 2] -> ... -> [path n].

3. The method according to claim 1, characterized in that In step S106, the UI element level file and the screenshot of the application page are preprocessed, specifically: Step S301: Traverse the UI element hierarchy file obtained in step S106, extract the UI elements whose clickable and focusable attribute values ​​are true, select the UI elements whose sizes do not exceed the threshold, extract the boundary coordinate information of these UI elements, and finally determine the numbering order of these UI elements according to the traversal order; Step S302: Based on the screenshot of the current application page obtained in step S106, and the UI elements and their numbering sequence and boundary coordinate information extracted in step S301, mark these controls with bounding boxes at the corresponding positions of each UI element on the screenshot, and mark their numbers directly above the bounding boxes.

4. The method according to claim 1, characterized in that: In step S107, the text information is generated for the UI element using the UI element hierarchy file, specifically: Step S401: Filter the text, content-desc, and resource-id attributes from the UI element hierarchy file. If all the three attributes are empty, proceed to step S402. Otherwise, concatenate the text information of the three attributes in order as the generated UI element text information. Step S402: Obtain a picture of the corresponding UI element from the screenshot according to the coordinates of the corresponding UI element obtained in step S301, input it into the multimodal large language model, and determine whether the UI element is a picture of displaying content or a UI element for switching functions. If it is the former, discard the UI element, otherwise proceed to step S403; Step S403: Obtain the image of the UI element according to step S402, combine it with the prompt word, give the functional description of the common application UI element, input it into the multimodal large language model, and let it output the text description corresponding to the UI element.

5. The method according to claim 1, characterized in that In step S108, the multimodal large language model outputs a driver program that converts the next application operation into a simulated operation, specifically: Step S501: The multimodal large language model structuredly outputs four types of operations, including click, input text, long click and slide, and the parameters used are respectively UI element number, input content, UI element number and UI element number plus direction and distance; Step S502: according to step S501, the operation output by the large language model is obtained. When a click operation is output, a click interface in the driver program of the simulated operation is called; When inputting or outputting text operations, you need to first call the click interface in the driver program of the simulation operation, and then call the text input interface; When a long tap and slide operation is output, the long tap and slide interface in the driver program of the simulated operation is called.

6. The method according to claim 1, characterized in that In step S108, whether a certain operation is successful and whether the application reaches the final path are determined, specifically: Step S601: As to whether a certain operation is successful, the application screenshots before and after the operation are input to the multimodal large language model. The large language model outputs three types of results according to the differences in the screenshots: successful operation, invalid operation and return. Successful operation means that a certain operation has achieved the expected effect, invalid operation means that there is no obvious change in the screenshots before and after the operation, and return means that the page entered by the current operation cannot meet the target page and needs to return to the previous page; Step S602: To determine whether the application has reached the final path, it is necessary to input a screenshot of the current application into the multimodal large language model, and combine it with prompt words to help the multimodal large language model output a determination result.

7. The method according to claim 1, characterized in that The step S110 determines the consistency result of the data path, specifically: Step S701: Input the text description of the UI element sequence obtained in S109 and the data path text obtained in step S102 into the word embedding model respectively, and output two vectors. Calculate the cross entropy loss value of the vectors. If it is greater than a threshold, the output data clicks are inconsistent, and if it is less than the threshold, the output data clicks are consistent; Step S702: Input the functional description of the final scene obtained in S108 and the functional description of the corresponding scene extracted in step S102 into the word embedding model respectively, and output two vectors. Calculate the cross entropy loss value of the vector. If it is greater than the threshold, the data path scene functional description is inconsistent. If it is less than the threshold, the output data path scene functional description is consistent.

8. According to step 701 and step 702 of claim 7, a threshold is used to determine whether the data click path is consistent with the scene function description. The threshold is calculated using a small sample with a label, and then extended to a large sample without a label, specifically: Step S801: manually marking the consistency of a small number of data paths, and after performing the operations of steps 101 to 109, calculating the average cross entropy loss value in the data click path samples to obtain a threshold for judging the consistency of the data click path; Step S802: Calculate the average cross entropy loss value of samples with consistent scene functions of the data path to obtain a threshold for determining the scene function consistency of the data path.