Mobile application automation script recording method and device, and storage medium
Patent Information
- Application Number
- CN202610981471.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-02
- Publication Date
- 2026-09-29
AI Technical Summary
[0004]本申请提供了一种移动应用自动化脚本录制方法、装置及存储介质,旨在解决目前由于无法直接捕获用户在手机端的自然操作行为,导致自动化测试脚本的录制结果准确性低的技术问题,以提高自动化脚本录制的录制结果准确性
[0008]本申请提供一种移动应用自动化脚本录制方法、装置及存储介质,本申请方法通过建立安卓调试桥通信连接并实时监听系统底层触控事件流,直接捕获用户在移动终端上的触控操作事件,避免了电脑端映射操作带来的行为偏差,保证了录制数据的真实性。基于预设的时间间隔和坐标偏移量判定条件对触控操作事件进行识别,确保操作类型判断的准确性;在识别操作类型的同时提取多维度的触控操作数据,为后续处理提供了完整的数字化描述基础。通过对比操作前后屏幕截图,判断当前触控操作的有效性,自动过滤误触、空白点击等无效操作,避免冗余步骤进入脚本,提高脚本录制效率和数据有效性。在确认操作有效后,基于预设的多级元素定位优先级策略确定元素定位方式,确保每个操作均使用最高优先级的可用定位策略,提高元素定位操作的效率和稳定性。根据元素定位方式执行元素定位操作,提取元素定位参数,并根据元素定位参数,将当前触控操作转换为结构化自然语言操作指令以完成脚本录制,使非技术人员能够直接理解和维护脚本内容,且保证了当前触控操作的可复现性,提高了自动化脚本录制结果的准确性。
Smart Images

Figure CN122838271A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of mobile application automated testing technology, and in particular to a method, apparatus and storage medium for recording automated scripts for mobile applications. Background Technology
[0002] In the field of mobile application (APP) automated testing, the recording and playback of automated test scripts is a core aspect of ensuring software quality. Currently, relevant APP automated testing solutions mainly include methods such as computer-side mapping operations and image recognition. However, these technologies require testers to operate on a computer or rely on a single visual recognition method during execution, failing to directly capture the user's natural operating behavior on the mobile device, resulting in low accuracy of the recorded automated test scripts.
[0003] Therefore, how to improve the accuracy of recording results of automated script recording is a technical problem that urgently needs to be solved in this field. Summary of the Invention
[0004] This application provides a method, apparatus, and storage medium for recording automated scripts for mobile applications, aiming to solve the technical problem that the recording results of automated test scripts are inaccurate due to the inability to directly capture the user's natural operating behavior on the mobile phone, so as to improve the accuracy of the recording results of automated scripts.
[0005] Firstly, this application provides a method for recording automated scripts for mobile applications, the method comprising the following steps: After establishing an ADB communication connection between the computer and the mobile terminal, the computer listens for touch operation events triggered by the user on the mobile terminal through the ADB log listening service. Based on at least one preset touch operation determination condition, the current touch operation corresponding to the touch operation event is identified, the operation type of the current touch operation is determined, and the touch operation data corresponding to the current touch operation is extracted. The current touch operation data includes the timestamp of the current touch operation, operation coordinates, the XML layout file of the current interface of the mobile terminal, and the current screen screenshot. The validity of the current touch operation is determined by comparing the current screenshot with the previous screenshot taken before the current touch operation occurred. When the current touch operation is determined to be a valid operation, the element positioning method of the current touch operation is determined based on a preset multi-level element positioning priority strategy; Based on the touch operation data, an element positioning operation is performed using the element positioning method to determine the element positioning parameters corresponding to the current touch operation; Based on the element positioning parameters, a structured natural language operation instruction corresponding to the current touch operation is generated to complete the automated script recording for the mobile application.
[0006] Secondly, this application also provides a mobile application automation script recording device, the mobile application automation script recording device comprising: The event listening module is used to listen for touch operation events triggered by the user on the mobile terminal through the ADB log listening service after the computer and the mobile terminal establish an ADB communication connection. The touch operation recognition module is used to identify the current touch operation corresponding to the touch operation event based on at least one preset touch operation judgment condition, determine the operation type of the current touch operation, and extract the touch operation data corresponding to the current touch operation. The current touch operation data includes the timestamp of the current touch operation, operation coordinates, the XML layout file of the current interface of the mobile terminal, and the current screen screenshot. The touch operation validity determination module is used to compare the current screenshot with the previous screenshot before the current touch operation occurred to determine the validity of the current touch operation; The element positioning method determination module is used to determine the element positioning method of the current touch operation based on a preset multi-level element positioning priority strategy when the current touch operation is determined to be a valid operation. The element positioning module is used to perform an element positioning operation based on the touch operation data and through the element positioning method to determine the element positioning parameters corresponding to the current touch operation. The structured data generation module is used to generate structured natural language operation instructions corresponding to the current touch operation based on the element positioning parameters, and to complete the automated script recording of the mobile application.
[0007] Thirdly, this application also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the steps of the mobile application automated script recording method described above.
[0008] This application provides a method, apparatus, and storage medium for recording automated scripts for mobile applications. The method establishes an Android debug bridge communication connection and monitors the system's underlying touch event stream in real time, directly capturing user touch operation events on the mobile terminal. This avoids behavioral deviations caused by mapping operations on the computer, ensuring the authenticity of the recorded data. Touch operation events are identified based on preset time intervals and coordinate offset judgment conditions, ensuring the accuracy of operation type determination. Simultaneously, multi-dimensional touch operation data is extracted, providing a complete digital description foundation for subsequent processing. By comparing screenshots before and after the operation, the validity of the current touch operation is determined, automatically filtering out invalid operations such as accidental touches and blank clicks, avoiding redundant steps in the script, and improving script recording efficiency and data validity. After confirming the operation's validity, the element positioning method is determined based on a preset multi-level element positioning priority strategy, ensuring that each operation uses the highest priority available positioning strategy, improving the efficiency and stability of element positioning operations. The system performs element positioning operations based on the element positioning method, extracts element positioning parameters, and converts the current touch operation into structured natural language operation instructions based on the element positioning parameters to complete script recording. This allows non-technical personnel to directly understand and maintain the script content, while ensuring the reproducibility of the current touch operation and improving the accuracy of automated script recording results. Attached Figure Description
[0009] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 This application provides a flowchart illustrating an embodiment of a mobile application automated script recording method. Figure 2 A flowchart illustrating the overall technical solution of a mobile application automated script recording method provided in this application embodiment; Figure 3 This is a schematic diagram of a dual-threaded parallel architecture provided in an embodiment of this application; Figure 4 A flowchart illustrating a multi-level element location priority strategy provided in this application embodiment; Figure 5 A schematic diagram of OCR text block matching provided in an embodiment of this application; Figure 6 This is a schematic diagram of the structure of a first embodiment of a mobile application automation script recording device provided in this application; Figure 7This is a schematic block diagram of the structure of a computer device provided in an embodiment of this application.
[0011] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0013] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the order described. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.
[0014] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0015] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of a mobile application automated script recording method provided in this application. The method is applied on a computer.
[0016] like Figure 1 As shown, the mobile application automated script recording method includes steps S101 to S106.
[0017] S101. After the computer establishes an ADB communication connection with the mobile terminal, it listens for touch operation events triggered by the user on the mobile terminal through the ADB log listening service.
[0018] In one embodiment, an ADB (Android Debug Bridge) communication connection is established between the computer and the mobile terminal to create a stable data transmission channel and capture user touch behavior data on the mobile terminal in real time. Specifically, the user can physically connect the mobile terminal to the computer via a Universal Serial Bus (USB) data cable, or establish a network connection by pre-configuring a wireless Android Debug Bridge mode. After the computer runs the recording program, it automatically calls the Android Debug Bridge toolset to detect the current connection status and sends a device authorization request to the mobile terminal. Once the user confirms the authorization on the mobile terminal, a bidirectional communication channel is formally established.
[0019] ADB (Android Debug Bridge) is a general-purpose debugging tool officially provided by the Android system, consisting of a client, a daemon process, and a server. The client runs on the computer, while the daemon process runs in the background of the mobile terminal, and the two communicate via a transmission protocol. This connection supports command-line instruction transmission, file transfer, port forwarding, and log stream push functions. In this solution, this connection simultaneously serves as both a control command issuance channel and a data acquisition and feedback channel, and is the only data interaction link between the computer and the mobile terminal.
[0020] like Figure 2 As shown, after the communication channel between the computer and the mobile terminal is established, the computer can use the ADB tool to obtain basic device information of the mobile terminal, including parameters such as device model, operating system version number, and screen physical resolution, which are used as reference parameters for subsequent element positioning calculations, coordinate conversion, and interface adaptation.
[0021] Simultaneously, an ADB log monitoring service is started on the computer, continuously capturing the system's underlying log stream from the mobile terminal using ADB log viewing commands. This log stream is system-level data output, containing multi-level information such as kernel events, application framework events, and hardware driver events. The recording program filters the log stream in real time, retaining only target event types related to interface interaction, including touch events, focus change events, window manager events, and activity lifecycle events.
[0022] Furthermore, the touch operation events include consecutive and adjacent press events and release events; the step of listening to the touch operation events triggered by the user on the mobile terminal through the ADB log listening service includes: listening to the system-level touch event stream of the mobile terminal through the ADB log listening service; when a press event that touches the screen is detected in the system-level touch event stream, marking the start time and starting coordinates of the current touch operation; continuously tracking the change of touch point coordinates after the press event until a release event that leaves the screen is detected, marking the end time and ending coordinates of the current touch operation.
[0023] After the recording program starts on the computer, it initializes the ADB log listening service module. This module sends a log subscription request to the daemon process on the mobile terminal by calling the log viewing interface provided by the ADB tool. Upon receiving the request, the daemon process on the mobile terminal pushes various event logs generated by the system's underlying layer to the computer in real time. The log data stream is transmitted line by line, with each line containing fields such as timestamp, process identifier, log level, log tag, and log content.
[0024] The recording program performs preliminary parsing of the raw log stream, identifying and extracting event log entries tagged with touch events, window manager, and activity manager, forming the system's underlying touch event stream. This system's underlying touch event stream refers to a continuous sequence of raw touch event log data output by the mobile terminal's operating system kernel and hardware abstraction layer, arranged in chronological order according to the mobile terminal's system clock. Each event record in the system's underlying touch event stream contains physical parameters such as event type, timestamp, touch point coordinates, touch pressure, and touch area, reflecting the user's actual touch operation behavior.
[0025] The recording program's event parsing thread scans the system's underlying touch event stream line by line. When it detects a press event identifier in the log content, it determines that the user's finger or stylus has made its first contact with the screen. At this point, the recording program extracts the system timestamp from this log entry as the start time of the current touch operation, and also extracts the horizontal and vertical coordinates of the touch point as the starting coordinates. The start time and starting coordinates are written into the data structure instance of the current touch operation.
[0026] The press event refers to the initial event generated when the user's finger or stylus first touches the screen surface. The change in electrical signal generated by the touch screen hardware is captured by the driver and processed by the system kernel. This event marks the beginning of a touch operation.
[0027] Between a press event and a release event, the system may generate several movement events, reflecting the sliding process of the touch point on the screen. Consecutive and adjacent press and release events together constitute a single touch operation event.
[0028] The lift-off event refers to the termination event generated after the touch screen hardware detects the disappearance of the capacitance or pressure signal when the user's finger or stylus leaves the screen surface. This event marks the end of a touch operation.
[0029] When the event parsing thread detects a lift-off event flag, it determines that the user's finger or stylus has left the screen, and the current touch operation ends. At this point, the system timestamp is extracted from the log entry as the termination time, and the touch point coordinates are extracted as the endpoint coordinates, and written into the data structure instance of the current touch operation.
[0030] This embodiment establishes an Android debugging bridge communication connection and monitors the system's underlying touch event stream in real time, achieving accurate capture of the user's natural touch operations. No mapping operation is required on the computer, and the recording experience is completely consistent with daily mobile phone use. At the same time, the raw data collection based on the system's underlying logs ensures the timing accuracy of coordinate collection and the authenticity of behavior, providing a high-fidelity data foundation for subsequent operation type identification, element positioning, and script generation. This effectively overcomes the technical defects of traditional solutions, such as fragmented operating environment and distorted recorded behavior.
[0031] S102. Based on at least one preset touch operation determination condition, identify the current touch operation corresponding to the touch operation event, determine the operation type of the current touch operation, and extract the touch operation data corresponding to the current touch operation. The current touch operation data includes the timestamp of the current touch operation, operation coordinates, the XML (Extensible Markup Language) layout file of the current interface of the mobile terminal, and the current screenshot.
[0032] The mobile application automated script recording method provided in this application is used to automatically record scripts for mobile applications (APPs) on mobile terminals. The mobile application (APP) refers to a mobile application that only supports three touch operations: click, long press, and swipe. Its user interface is built based on the standard view architecture of the Android system and can generate XML layout files through the interface automation testing framework provided by the Android system. The operation types recognized by the recording program are limited to click, long press, and swipe, excluding complex multi-touch operations such as two-finger zoom, rotation, and drag, and also excluding interface types that rely entirely on custom rendering engines and cannot output a standard view tree structure.
[0033] In one embodiment, the start time, end time, start coordinates, and end coordinates of the current touch operation are obtained; the time interval between the start time and the end time is calculated, and the coordinate offsets of the start coordinates and the end coordinates are calculated; when the time interval and the coordinate offsets simultaneously satisfy any preset touch operation determination condition, the operation type of the current touch operation is determined.
[0034] After receiving the touch operation data corresponding to the current touch operation, the recording program extracts four basic parameters from the touch operation data: start time, end time, start coordinates, and end coordinates.
[0035] The start time is the system timestamp marked when the press event occurs, recording the precise time when the user's finger or stylus first touches the screen. The end time is the system timestamp marked when the release event occurs, recording the precise time when the user's finger or stylus leaves the screen. The start coordinates are the horizontal and vertical coordinates of the screen touch point corresponding to the start time, in pixels, with the origin of the coordinate system located at the top left corner of the screen, increasing horizontally to the right and vertically downwards. The end coordinates are the horizontal and vertical coordinates of the screen touch point corresponding to the end time, using the same coordinate system as the start coordinates.
[0036] The time interval refers to the time difference between the start moment when a finger or stylus touches the screen and the end moment when it leaves the screen during a single touch operation. The time interval can be calculated by taking the difference between the start and end moments. It reflects the duration of a user's press and is a core time dimension characteristic that distinguishes between clicks, long presses, and swipes.
[0037] Coordinate offset refers to the spatial distance between the starting and ending coordinates in a single touch operation. It can be calculated using the Euclidean distance formula for a two-dimensional plane, specifically the square root of the sum of the squares of the differences in the x-coordinates and y-coordinates, yielding a distance value in pixels as the coordinate offset. This coordinate offset reflects the magnitude of finger displacement during the user's operation and is a core spatial dimension feature distinguishing between click, long press, and swipe operations. The coordinate offsets for click and long press operations are typically caused by slight finger jitter or measurement errors when the finger touches the screen, and are relatively small. The coordinate offset for swipe operations reflects the distance the user actively swipes, is larger, and has a clear direction. Furthermore, for swipe operations, the coordinate offset also implicitly contains swipe direction information, which can be determined by the sign and relative magnitude of the differences in the x-coordinate and y-coordinates.
[0038] The time interval and coordinate offset are used as inputs for judgment, and are compared one by one with the judgment conditions for three preset touch operations (click, long press, and swipe). The judgment conditions for the preset touch operations refer to a set of operation type recognition rules predefined and stored in the recording program configuration. The judgment conditions can be set based on the principles of human-computer interaction engineering, reflecting the natural temporal and spatial scales of human finger operations.
[0039] Because the mobile application targeted by this application only supports three basic touch operations: click, long press, and swipe, and these three operations have significantly different spatiotemporal distribution patterns in terms of human-computer interaction characteristics, the click operation is characterized by a quick and short single-point touch, the long press operation is characterized by a continuous and stable press and hold, and the swipe operation is characterized by a directional displacement of a certain duration, this application sets different judgment condition threshold ranges for these three preset touch operations to accurately identify the user's true operation intention.
[0040] In one embodiment, when the time interval is within a first interval range and the coordinate offset is within a first offset distance range, the operation type of the current touch operation is determined to be a click operation.
[0041] This application sets a first interval range and a first offset distance range as the criteria for determining a click operation. The first interval range is a time threshold interval within the criteria for determining a click operation. This first interval range defines the typical time window during which a user quickly touches the screen and immediately leaves, reflecting the short and quick interaction characteristics of a click operation. For example, the first interval range can be set to 0 milliseconds to 300 milliseconds, covering the typical time characteristics of a user quickly touching the screen and immediately leaving; the first offset distance range can be set to 0 pixels to 10 pixels, covering the slight displacement caused by minor finger tremors or touchscreen sampling errors during the click.
[0042] The calculated time interval is compared with the first interval range. If the time interval is within the first interval range, the time interval is determined to meet the time condition for the click operation. At the same time, the calculated coordinate offset is compared with the first offset distance range. If the coordinate offset is within the first offset distance range, the coordinate offset is determined to meet the spatial condition for the click operation.
[0043] When both the time interval and the coordinate offset satisfy their respective range conditions, i.e. the time interval is within the first interval range and the coordinate offset is within the first offset distance range, it indicates that the user is performing a fast, short-distance single-point touch action, which is consistent with the human-computer interaction characteristics of a click operation. Therefore, the operation type of the current touch operation is determined to be a click operation.
[0044] For example, if the first interval range is set to [0, 300] milliseconds, and the first offset distance range is set to [0, 10] pixels, the tester clicks the login button on the homepage of the application under test. The recording program obtains the current touch operation, reading the start time as 14:20:10:150 and the end time as 14:20:320, the starting coordinates as 400 pixels x and 800 pixels y, and the ending coordinates as 402 pixels x and 798 pixels y. The calculated time interval is 170 ms, and the coordinate offset is 2.8 pixels. The time interval of 170 ms is within the first interval range of [0, 300] milliseconds, and the coordinate offset of 2.8 pixels is within the first offset distance range of [0, 10] pixels. Both parameters simultaneously satisfy the click operation judgment condition, determining that the current touch operation type is a click operation.
[0045] In one embodiment, when the time interval is within a second interval range and the coordinate offset is within a second offset distance range, the operation type of the current touch operation is determined to be a long press operation.
[0046] This application sets a second interval range and a second offset distance range as the criteria for determining a long-press operation. A long-press operation requires the user's finger or stylus to continuously press at the same position on the screen for a period of time. Therefore, the second interval range needs to be set to a relatively long time range, such as [500, 3000], in milliseconds, covering the typical time characteristics of a user continuously pressing the screen to trigger a long-press response. Simultaneously, since the user's finger may slightly slide due to muscle fatigue during a long press, the second offset distance range can allow for a slightly larger finger displacement margin than a click operation, such as setting the second interval range to [0, 15], in pixels. When both the time interval and coordinate offset fall within the above two ranges, it indicates that the user is performing a continuous pressing action that remains essentially still, consistent with the human-computer interaction characteristics of a long-press operation.
[0047] If the combination of time interval and coordinate offset does not meet the criteria for a click operation, the time interval is compared with the second interval range. If the time interval is within the second interval range, the time interval is determined to meet the time condition for a long press operation. Simultaneously, the coordinate offset is compared with the second offset distance range. If the coordinate offset is within the second offset distance range, the coordinate offset is determined to meet the spatial condition for a long press operation. When both the time interval and coordinate offset parameters simultaneously meet the criteria for a long press operation (i.e., the time interval is within the second interval range and the coordinate offset is within the second offset distance range), the current touch operation is determined to be a long press operation.
[0048] For example, if the second interval range is set to [500, 3000] in milliseconds (ms), and the second offset distance range is set to [0, 15] in pixels, a tester long-presses an image in the application under test to trigger the delete option. The recording program calculates the time interval to be 2800 milliseconds and the coordinate offset to be 1.4 pixels. The click operation is then evaluated: the time interval of 2800 milliseconds exceeds the upper limit of the first interval range by 300 milliseconds, thus not meeting the click operation condition. The long-press operation is then evaluated: the time interval of 2800 milliseconds is greater than or equal to 500 milliseconds and less than or equal to 3000 milliseconds, falling within the second interval range; the coordinate offset of 1.4 pixels is greater than 0 pixels and less than 15 pixels, falling within the second offset distance range. Since both the time interval and the coordinate offset meet the criteria for a long-press operation, the current touch operation is determined to be a long-press operation.
[0049] In one embodiment, when the time interval is within a third interval range and the coordinate offset is within a third offset distance range, the operation type of the current touch operation is determined to be a swipe operation.
[0050] This application sets a third interval range and a third offset distance range as the criteria for determining a swipe operation. The third interval range defines the typical time window for a user to actively swipe the screen to browse content, reflecting the interactive characteristic that a swipe operation requires a certain duration to complete the displacement. For example, the third interval range can be set to [100, 2000] milliseconds, covering the typical time characteristics of a user actively swiping the screen to browse content. The third offset distance range defines the obvious displacement that a swipe operation must produce, reflecting the spatial characteristic of a swipe operation with directional movement as its core. For example, the third offset distance range can be set to [30, screen diagonal length] pixels. The lower limit is set to 30 pixels to clearly distinguish it from the small displacements of click and long press operations, eliminating misjudgments caused by slight jitter during clicks or long presses. The upper limit is set to the screen diagonal length to cover swipe scenarios of any direction and any distance within the entire screen range, including the extreme case of swiping from one corner of the screen to the opposite corner.
[0051] When the time interval and coordinate offset both satisfy the third interval range and the third offset distance range, it indicates that the user is performing a swipe gesture with a certain duration and significant displacement, which conforms to the human-computer interaction characteristics of swipe operation.
[0052] If the combination of time interval and coordinate offset does not meet either the click operation or long press operation criteria, the recording program continues to execute the swipe operation criteria. The time interval is compared with the third interval range; if the time interval falls within the third interval range, the time interval is determined to meet the time condition for a swipe operation. Simultaneously, the coordinate offset is compared with the third offset distance range; if the coordinate offset falls within the third offset distance range, the coordinate offset is determined to meet the spatial condition for a swipe operation. When both the time interval and coordinate offset parameters simultaneously meet their respective range conditions, the recording program determines the current touch operation type as a swipe operation.
[0053] For example, if the third interval range is set to [100, 2000] in milliseconds (ms), and the third offset distance range is set to [30, 1000] in pixels, where the upper limit of 1000 represents the screen diagonal length, and the tester scrolls down through the application list, the recording program calculates a time interval of 600 milliseconds and a coordinate offset of 800.6 pixels. The 600 millisecond time interval falls within [100, 2000], satisfying the third interval range; the 800.6 pixel coordinate offset falls within [30, 1000], satisfying the third offset distance range. Both parameters simultaneously meet the swipe operation determination criteria, confirming the current touch operation type as a swipe operation.
[0054] If the combination of time interval and coordinate offset does not meet any of the three judgment conditions mentioned above, the current touch operation is judged as an invalid touch operation. The recording program terminates subsequent processing of this operation, records the filtering log, and waits for the next touch operation event to be triggered. It can be understood that the threshold range corresponding to the judgment conditions of each preset touch operation can be dynamically adjusted through the configuration file to adapt to the operating habits of different user groups or the needs of specific test scenarios.
[0055] This embodiment achieves accurate identification and classification of three basic touch operations—click, long press, and swipe—through a differentiated threshold determination mechanism based on spatiotemporal dual-dimensional features (time interval and coordinate offset). This simplifies the complexity of the identification logic and improves the accuracy and processing efficiency. Simultaneously, the dual threshold verification mechanism using time interval and coordinate offset effectively filters out invalid or abnormal operations such as accidental touches and jitter, ensuring the purity of the recorded script's operations and the reliability of its playback. This provides a high-quality data foundation for subsequent element positioning and script generation.
[0056] After determining the type of the current touch operation, the recording program executes the touch operation data extraction process. Specifically, it reads the start and end times as the timestamps of the current touch operation; and reads the start coordinates as the operation coordinates of the current touch operation. Then, the recording program sends an interface layout capture command to the mobile terminal via the ADB channel. After receiving the command, the mobile terminal's interface automation testing framework traverses the view tree structure of the currently active window, generates an XML layout file containing the attribute information of all user interface nodes, and sends it back to the computer. Simultaneously, the recording program sends a screen capture command to the mobile terminal via the ADB channel, captures a screenshot of the current screen of the mobile terminal, and sends it back to the computer.
[0057] The XML layout file refers to the view tree structure description file generated by the Android system's UI automation testing framework. Encoded in Extensible Markup Language (XML) format, this file contains attribute information for all visible and invisible user interface nodes in the currently active window. Each node corresponds to a UI element, and its attribute information includes structured data such as text content, resource identifiers, content descriptions, node types, and boundary coordinates. This file is a standard UI structure data source provided by the Android system, offering precise structured information for element location.
[0058] The current screenshot refers to a complete pixel data snapshot of the mobile terminal screen at the time of the current touch operation. The screenshot content includes all visual elements of the current interface, including text, icons, images, animation frames, etc., and is used for image comparison and analysis in subsequent valid operation judgment.
[0059] The timestamp, operation coordinates, XML layout file, current screenshot, and determined operation type are encapsulated into touch operation data corresponding to the current touch operation. Touch operation data refers to a multi-dimensional data set collected around a single touch operation. This data set includes time dimension information (timestamp), spatial dimension information (operation coordinates), interface structure information (Extensible Markup Language layout file), and interface visual information (current screenshot), forming a complete digital description of the touch operation.
[0060] S103. Compare the current screenshot with the previous screenshot before the current touch operation to determine the validity of the current touch operation.
[0061] Please refer to Figure 3 , Figure 3 This is a schematic diagram of a dual-threaded parallel architecture provided in an embodiment of this application. Figure 3 As shown, the embodiments of this application adopt a dual-thread parallel architecture, which separates data monitoring and acquisition and script recording and processing into two independently running threads to achieve asynchronous parallel execution of data acquisition and data processing.
[0062] Specifically, thread A is a data listening and acquisition thread responsible for executing ADB log listening services and capturing touch operation events. After establishing a log subscription channel in the mobile terminal's daemon process, thread A continuously receives the system's underlying touch event stream, performs real-time parsing of the touch event stream, identifies press and release events, and encapsulates the collected touch operation data into a data object after each complete touch operation event (i.e., from press to release). Simultaneously, thread A is also responsible for sending interface layout capture and screen capture commands to the mobile terminal via the ADB channel after each touch operation, obtaining the current interface's XML layout file and the current screenshot, and encapsulating information such as the timestamp, operation coordinates, XML layout file, current screenshot, and operation type into touch operation data.
[0063] Thread B is the script recording and processing thread, responsible for performing tasks such as validating the current touch operation, determining the element positioning method, executing the element positioning operation, and generating structured natural language operation instructions. After receiving touch operation data from thread A, thread B sequentially executes the processing flow, including validating the operation, determining the multi-level element positioning priority strategy, extracting element positioning parameters, and generating structured natural language operation instructions.
[0064] Thread A and Thread B interact via a thread-safe data transmission channel. Specifically, Thread A writes the encapsulated touch operation data to a shared buffer or message queue, and Thread B reads the touch operation data from the shared buffer or message queue for processing. The thread-safe queue acts as the data exchange hub between Thread A and Thread B, ensuring the security and consistency of data transmission in a multi-threaded environment. Data packets generated by Thread A are enqueued in chronological order, and Thread B consumes and processes them in the same order, avoiding data out-of-order processing or loss. The structured natural language operation instructions generated by Thread B are first written to a memory buffer for temporary caching during recording. When the user actively ends recording or the preset automatic save conditions are met, all operation instructions in the memory buffer are batch-written to persistent storage, completing the final generation of the script file.
[0065] This embodiment uses a dual-thread parallel architecture to achieve asynchronous parallel execution of data acquisition and data processing. Thread A continuously monitors and acquires data without being affected by the processing time of thread B, and thread B processes data independently without affecting the real-time monitoring of thread A. This effectively improves the throughput and real-time performance of the recording system and avoids the operation omissions or delays caused by mutual blocking between data acquisition and data processing in a single-thread architecture.
[0066] Specifically, such as Figure 2 and Figure 3 As shown, after determining the operation type of the current touch operation, the interface changes before and after the operation are compared using image recognition technology to determine whether the user operation is valid and to automatically filter out invalid operations in order to improve the quality of the recorded script.
[0067] First, the current screenshot is extracted from the touch operation data. Simultaneously, the previous screenshot before the current touch operation is retrieved from the memory buffer or persistent storage. The previous screenshot is the interface image captured when the last valid touch operation was completed. If the current touch operation is the first operation during recording, the previous screenshot defaults to the initial interface screenshot taken when recording started.
[0068] After acquiring the current and previous screenshots, a difference analysis is performed on the two screenshots. For example, a pixel-level comparison algorithm is used to compare the two screenshots of the same size pixel by pixel, calculating the number of pixels whose color values differ by more than a preset tolerance threshold. After counting the number of differing pixels, the proportion of these differing pixels to the total number of pixels in the screenshot is calculated as a quantitative indicator of the degree of interface change. The tolerance threshold is set to a reasonable range that takes into account minor fluctuations in screen display, avoiding misjudgments caused by non-operational factors such as system status bar icon refreshes and changes in loading animation frames.
[0069] The calculated degree of interface change is compared with a preset validity threshold to determine whether the current touch operation is valid or invalid. Generally, a valid operation is a touch operation that causes a significant change in the interface, that is, an operation whose degree of interface change reaches or exceeds the validity threshold. This usually corresponds to interactive behaviors with clear user intent, such as clicking a button to trigger a page jump, long-pressing an element to pop up a menu, or sliding a list to scroll content. An invalid operation is a touch operation that does not cause a visible change in the interface, that is, an operation whose degree of interface change is below the validity threshold. This is usually caused by accidental touches, clicking on blank areas, repeatedly clicking on selected elements, or clicking on disabled buttons.
[0070] Specifically, if the degree of interface change is below the validity threshold, it means that the current touch operation did not cause a visible change in the interface. This could be due to accidental touches, clicking on blank areas, or repeatedly clicking on already selected elements. In this case, the operation is determined to be invalid, the recording program automatically filters it, records the filtering log, terminates the subsequent element location and script generation process, and waits for the next touch operation event to be triggered. If the degree of interface change reaches or exceeds the preset threshold, it means that the current touch operation caused a significant change in the interface, and the operation is determined to be valid. The recording program continues to execute the subsequent element location steps.
[0071] This embodiment analyzes the pixel-level differences between screenshots taken before and after an operation, automatically filtering out invalid operations such as accidental touches, blank clicks, and repeated operations, thus avoiding redundant steps from entering the recording script and significantly improving the script's usability and playback success rate.
[0072] S104. When the current touch operation is determined to be a valid operation, the element positioning method of the current touch operation is determined based on a preset multi-level element positioning priority strategy.
[0073] like Figure 2 , Figure 3 and Figure 4 As shown, after determining that the current touch operation is valid, the element positioning method for the current touch operation is selected according to a preset multi-level element positioning priority strategy, and then the element positioning operation is performed using the selected element positioning method. The multi-level element positioning priority strategy contains three levels, which are judged sequentially in a fixed order. The element positioning method of each level is only enabled when the previous level is determined to be unavailable. That is, the availability of each level of positioning method is judged in a fixed order. Once a level is determined to be available, subsequent judgments are immediately stopped, and the positioning method of that level is determined as the element positioning method for the current touch operation. This ensures that each valid operation uses the highest priority available positioning strategy, avoids the overriding of high-priority strategies by low-priority strategies, and guarantees the efficiency of element positioning operations as well as the accuracy and stability of positioning results.
[0074] In one embodiment, the multi-level element positioning priority strategy determines the positioning methods at each level in the order of first priority positioning method, second priority positioning method, and third priority positioning method, and determines the first positioning method that is determined to be available as the element positioning method of the current touch operation.
[0075] Please refer to Figure 4 , Figure 4 A flowchart illustrating the multi-level element location priority strategy provided in this application embodiment. For example... Figure 4 As shown, the multi-level element location priority strategy includes three levels that are judged sequentially: the first priority location method (XML attribute location method), the second priority location method (OCR text location method), and the third priority location method (edge detection location method).
[0076] After determining that the current touch operation is valid, the recording program first enters the first priority determination stage, performing an availability determination of the XML attribute positioning method. This involves searching for the corresponding target user interface node in the current interface's XML layout file based on the operation coordinates in the touch operation data, and checking whether the node contains valid attribute information. If the XML attribute positioning method is determined to be available, it is directly identified as the element positioning method for the current touch operation, the process ends, and no further level of determination is performed.
[0077] If the XML attribute positioning method is determined to be unavailable, the process proceeds to the second priority determination stage, where the availability of the OCR text positioning method is determined. This involves performing optical character recognition (OCR) on the current screenshot to determine if the target text content exists at the operation coordinates. If the OCR text positioning method is determined to be available, it is then designated as the element positioning method for the current touch operation, and the process ends.
[0078] If the OCR text positioning method is determined to be unavailable, the process continues to the third priority determination stage, where the availability of the edge detection positioning method is determined. This involves performing edge detection on the current screenshot to identify control icons and matching them using the local model. If the edge detection positioning method is determined to be available, it is selected as the element positioning method for the current touch operation, and the process ends.
[0079] If all three positioning methods are deemed unavailable, an external image understanding model is invoked for fallback identification, or the current touch operation is marked as an abnormal operation. Therefore, Figure 4 The multi-level element location priority strategy shown is determined step by step in a fixed order. Each level is only activated when the previous level is unavailable. Once a level is determined to be available, subsequent determinations are stopped immediately to ensure that each valid operation uses the highest priority available location strategy.
[0080] The following is a detailed description of the specific determination process for each priority element location method in the multi-level element location priority strategy: In one embodiment, the first priority location method is the XML attribute location method.
[0081] When determining the first priority positioning method, based on the operation coordinates in the touch operation data, the corresponding target user interface node is searched in the XML layout file of the current interface; it is detected whether the target user interface node contains valid attribute information, wherein the attribute information includes at least one of identity attribute, text attribute and content description attribute; when the target user interface node contains valid attribute information, it is determined that the first priority positioning method is available, and the element positioning method of the current touch operation is determined to be the XML attribute positioning method.
[0082] First, read the operation coordinates from the touch operation data corresponding to the current touch operation. These operation coordinates are the starting coordinates of the current touch operation, represented in the screen pixel coordinate system, with the horizontal and vertical coordinate values accurate to a single pixel.
[0083] Based on the operation coordinates, the corresponding target user interface node is searched within the current interface's XML layout file. The search process can employ a range traversal strategy centered on the operation coordinates and with a preset search radius as the spatial boundary. Specifically, the document object model (DOM) tree structure of the Extensible Markup Language (XML) layout file is parsed to obtain the boundary coordinate attributes of all user interface nodes. These boundary coordinate attributes are defined as rectangular regions containing four values: the x-coordinate of the top-left corner, the y-coordinate of the top-left corner, the x-coordinate of the bottom-right corner, and the y-coordinate of the bottom-right corner. Each user interface node's boundary coordinate attribute is checked sequentially to see if it contains the operation coordinates. Specifically, the x-coordinate value is determined to be greater than or equal to the top-left x-coordinate and less than or equal to the bottom-right x-coordinate, and the y-coordinate value is also determined to be greater than or equal to the top-left y-coordinate and less than or equal to the bottom-right y-coordinate. If multiple user interface nodes contain the operation coordinates in their boundary coordinate attributes, a score is awarded based on factors such as the node's hierarchy depth in the DOM tree and its area size, and the optimal node is selected as the target user interface node.
[0084] The preset search radius can be set to fifty pixels. When a user interface node containing operation coordinates is not located within the default search radius, the recording program supports dynamically expanding the search radius according to the screen resolution and re-executing the traversal search until the target user interface node is located or the maximum search radius limit is reached.
[0085] After locating the target user interface node, the recording program checks whether the target user interface node contains valid attribute information. Valid attribute information includes at least one of the following: identity attribute, text attribute, and content description attribute. The identity attribute corresponds to the identity field of user interface elements in the Android system, which is a unique identifier string explicitly defined by the developer in the layout file or code; the text attribute corresponds to the visible text content displayed by the user interface element, such as text labels on buttons, prompt text in input boxes, etc.; the content description attribute corresponds to the content description field in the Android system's accessibility services.
[0086] During the detection process, the recording program sequentially reads the identity attribute value, text attribute value, and content description attribute value of the target user interface node, and performs validity checks. The validity check rules are: the attribute value must not be an empty string, must not be an empty value, must not be a meaningless string containing only whitespace characters, and must not be a system-generated random placeholder. If at least one of the identity attribute, text attribute, and content description attribute passes the validity check, the target user interface node is determined to contain valid attribute information.
[0087] When the target user interface node contains valid attribute information, the first priority positioning method is determined to be available, and the element positioning method for the current touch operation is determined to be XML attribute positioning. At this time, the determination of the second and third priority positioning methods is no longer performed, and the element positioning operation is directly performed using the XML attribute positioning method.
[0088] This embodiment uses XML attribute positioning as the first priority positioning method, which can directly use the structured interface data provided by the Android system to locate elements without the need for complex image recognition or visual analysis calculations. The positioning speed is fast and the results are accurate and stable.
[0089] If the target user interface node does not contain any valid attribute information, that is, the identity attribute, text attribute, and content description attribute are all empty or have failed the validity check, then the first priority positioning method is deemed unavailable. In this case, the second priority positioning method needs to be determined.
[0090] In one embodiment, the second priority positioning method is an OCR (Optical Character Recognition) text positioning method. The step of determining the element positioning method of the current touch operation based on a preset multi-level element positioning priority strategy includes: when the first priority positioning method is unavailable, performing OCR text recognition processing on the current screenshot to obtain the text content in the current screenshot and the location information corresponding to the text content; based on the operation coordinates of the current touch operation and the location information, determining whether there is corresponding target text content at the operation coordinate location; when the target text content exists at the operation coordinate location, determining that the second priority positioning method is available, and determining that the element positioning method of the current touch operation is the OCR text positioning method.
[0091] When determining the OCR text localization method, the current screenshot is obtained from the touch operation data corresponding to the current touch operation. Then, the optical character recognition engine is called to perform text recognition processing on the current screenshot. Specifically, the optical character recognition engine uses a deep learning-based text detection and recognition algorithm. First, the input image is preprocessed, including grayscale conversion, binarization, noise removal, and tilt correction, to improve recognition accuracy. Then, the engine uses a text region detection network to locate the positions of text blocks in the image and outputs the bounding rectangle coordinates of each text block. Next, the engine decodes the character sequence within each text block using a text recognition network and outputs the recognized text content. Finally, the optical character recognition engine outputs all the text content in the current screenshot, along with the position information corresponding to each segment of text. The position information can be represented as a rectangular region, containing four values: the top-left horizontal coordinate, the top-left vertical coordinate, the bottom-right horizontal coordinate, and the bottom-right vertical coordinate, accurately describing the spatial range occupied by the text segment in the screenshot.
[0092] Then, the operation coordinates of the current touch operation, i.e., the starting coordinates of the current touch operation, are obtained, and these operation coordinates are compared one by one with the position information of each text block in the optical character recognition result. The comparison method is as follows: it is determined whether the horizontal coordinate value of the operation coordinate is greater than or equal to the horizontal coordinate of the upper left corner of a text block position information, and less than or equal to the horizontal coordinate of the lower right corner; at the same time, it is determined whether the vertical coordinate value of the operation coordinate is greater than or equal to the vertical coordinate of the upper left corner of the text block position information, and less than or equal to the vertical coordinate of the lower right corner. If the operation coordinate satisfies both of the above conditions, it indicates that the operation coordinate falls within the rectangular area defined by the text block position information, and the operation coordinate position contains text content corresponding to the text block.
[0093] Since optical character recognition results may contain multiple text blocks, and these text blocks may overlap or be nested, if the operation coordinates fall within the rectangular area of a single text block's location information, the text content of that text block is directly identified as the target text content. If the operation coordinates fall within the rectangular area of multiple text block locations, the text blocks are sorted from smallest to largest, and the smallest text block is selected first.
[0094] For example, in Figure 5 In the illustrated OCR text block matching diagram, the optical character recognition engine performs text recognition on the current screenshot and outputs three text blocks: text block A, text block B, and text block C. Text block A has the largest rectangular area, text block B's rectangular area is nested inside text block A and is centered, and text block C's rectangular area is further nested inside text block B and is the smallest, forming a nested relationship from the outside in. The operation coordinate P falls within the rectangular areas of text blocks A, B, and C simultaneously. Following a strategy of sorting by area from smallest to largest, the recording program compares the areas of the three text blocks and determines that text block C has the smallest area; therefore, the text content of text block C is identified as the target text content. The principle behind this minimum area priority strategy is that smaller text blocks correspond to finer-grained interface elements, resulting in higher matching accuracy with the user's actual operation target and effectively avoiding positioning deviations caused by selecting outer, coarse-grained text blocks.
[0095] If the operation coordinates do not fall within the rectangular area of any text block location information, but are within the first adjacent range of a rectangular area of a text block location information, then the distance between the operation coordinates and the center point of that rectangular area is calculated, and the nearest text block is selected as a candidate. The first adjacent range can be set according to the required accuracy; for example, setting the first adjacent range to 10 pixels can accommodate position detection errors from the optical character recognition engine and slight deviations in user clicks.
[0096] If the operation coordinates do not fall within the rectangular area of any text block position information, nor within the first adjacent range, it is determined that there is no corresponding target text content at the operation coordinate position.
[0097] After the above judgment process, if the target text content is detected at the position corresponding to the operation coordinates in the current screenshot, it is determined that the third priority positioning method is available. At this time, the element positioning method of the current touch operation is determined to be the OCR text positioning method.
[0098] In this embodiment, optical character recognition text positioning is set as the second priority. When the structured attribute positioning of the first priority fails, visual text recognition is performed on the screenshot and the operation coordinate position is matched. This does not rely on the structured data provided by the application's underlying code or system framework. Therefore, it can effectively cover scenarios that the standard user interface tree cannot fully reflect, such as web page view pages, canvas self-drawn interfaces, and dynamically loaded content. It fills the positioning blind spot of the first priority under special interface types and improves the compatibility and adaptability of script recording in complex interface scenarios.
[0099] If, after the above judgment logic, the recording program determines that there is no corresponding target text content at the operation coordinate position, that is, the optical character recognition engine did not recognize any text block, or the recognized text block positions do not intersect with the operation coordinates, or the text block cannot be accurately recognized due to blurring or occlusion, then the second priority positioning method is deemed unusable. At this time, the recording program continues to execute the judgment of the third priority positioning method.
[0100] In one embodiment, the third priority positioning method is an edge detection positioning method. The step of determining the element positioning method of the current touch operation based on a preset multi-level element positioning priority strategy further includes: when the second priority positioning method is unavailable, performing edge detection on the current screenshot to identify control icons in the current screenshot and their corresponding position information; determining a target control icon based on the operation coordinates of the current touch operation and the position information of the control icon; inputting the target control icon into a local model for recognition and obtaining the recognition result of the local model; when the recognition result indicates that the target control icon has been successfully recognized, determining that the third priority positioning method is available, and determining that the element positioning method of the current touch operation is an edge detection positioning method.
[0101] Get the current screenshot from the touch operation data corresponding to the current touch operation, and call the edge detection algorithm to perform control icon recognition processing on the current screenshot.
[0102] Specifically, the edge detection algorithm is based on edge detection operators in computer vision. First, the input image is converted to grayscale, transforming a color image into a single-channel grayscale image to reduce subsequent computational complexity. Then, Gaussian filtering is used to smooth the image, suppressing noise interference and preventing noise points from being misidentified as edges. Next, the gradient intensity and gradient direction of each pixel in the image are calculated. The gradient intensity reflects the degree of grayscale change in the neighborhood of that pixel, and the gradient direction reflects the direction of the grayscale change. Pixels with gradient intensities greater than a preset intensity threshold are marked as edge points, forming a preliminary edge image.
[0103] Based on the initial edge image, the edge detection algorithm refines the edges through non-maximum suppression, retaining only local maxima points along the gradient direction and eliminating edge width to obtain precise edges with a single pixel width. Then, using a dual-threshold detection and edge connection strategy, edge points are classified into three categories: strong edges, weak edges, and non-edges. Strong edge points are directly identified as the final edges, weak edge points connected to strong edge points are also retained, and isolated weak edge points are removed, forming a continuous and complete edge contour.
[0104] Based on the detected edge contours, regions with regular geometric shapes are identified as candidate control icons through contour analysis and ensemble feature extraction. Control icons typically have features such as closed contour boundaries, regular rectangular or circular shapes, and color contrast that is clearly distinguishable from the background. Therefore, the edge detection algorithm can calculate feature vectors such as contour area, aspect ratio, roundness, and color consistency for each candidate control icon, and match them with a preset control icon feature template to filter out control icon regions with high confidence.
[0105] After completing feature matching for all candidate control icons, the edge detection algorithm outputs a list of identified control icons in the current screenshot, along with the location information for each icon. The location information is represented as a rectangular area, containing four values: the x-coordinate of the top-left corner, the y-coordinate of the top-left corner, the x-coordinate of the bottom-right corner, and the y-coordinate of the bottom-right corner, precisely describing the spatial extent of the control icon in the screenshot. Simultaneously, the edge detection algorithm outputs an edge feature description for each control icon, including its outline shape, dominant color, and texture features, serving as auxiliary input for subsequent local model recognition.
[0106] After edge detection is completed, the position information of each control icon in the control icon list output by the edge detection algorithm is compared one by one with the operation coordinates of the current touch operation. The comparison method can be described as follows: determine whether the horizontal coordinate value of the operation coordinate is greater than or equal to the horizontal coordinate of the top left corner and less than or equal to the horizontal coordinate of the bottom right corner of the control icon's position information, and at the same time, whether the vertical coordinate value of the operation coordinate is greater than or equal to the vertical coordinate of the top left corner and less than or equal to the vertical coordinate of the bottom right corner of the control icon's position information. If the operation coordinate satisfies both of the above conditions, it means that the operation coordinate falls within the rectangular area defined by the control icon's position information, and the control icon is identified as the target control icon clicked by the user.
[0107] If the operation coordinates do not fall within the rectangular area of any control icon's position information, but are within the second adjacent range of a rectangular area of a control icon's position information, then the distance between the operation coordinates and the center point of that rectangular area is calculated, and the control icon closest to it is selected as the target control icon. The second adjacent range can be set to 15 pixels to accommodate position detection errors from the edge detection algorithm and slight deviations in user clicks.
[0108] If the operation coordinates do not fall within the rectangular area containing any control icon position information, nor within the second adjacent range, it is determined that the target control icon cannot be identified, and the third priority edge detection positioning method fails.
[0109] In one embodiment, after determining the target control icon, the target control icon is cropped and extracted from the current screenshot to form an independent icon image. The cropping area can be a rectangular area based on the position information of the target control icon, with a preset margin (e.g., 5 pixels) extended outward to ensure icon integrity.
[0110] The cropped icon image is input into a local model for recognition. The local model can be a lightweight image classification model deployed on a computer, built on a deep learning framework, primarily consisting of a feature extraction network and a classifier network. Simultaneously, a corresponding feature library can be configured for the local model to store the feature vectors and identifier names of registered spatial icons (such as the icons of all apps on a mobile device). The feature vectors are extracted by the model from known icon samples during the training or registration phase, and the identifier names are human-readable descriptive strings, such as "Home," "Search," "Settings," and "Personal Center."
[0111] When the icon image corresponding to the current touch operation is input into the local model, the feature extraction network in the local model extracts high-dimensional abstract feature vectors from the input icon image through multi-layer convolution and pooling operations, capturing visual features such as the shape, texture, color, and structure of the icon image. Then, a classifier network maps the feature vectors to the registered control icon category space. Using feature matching algorithms (such as cosine similarity or Euclidean distance), the similarity between the feature vector of the input icon image and the feature vectors of each registered icon in the feature library is calculated. Finally, the local model outputs a ranking list of the similarity scores between the image icon and each registered icon in the feature library, serving as the recognition result of the local model. This recognition result includes the identifier name and similarity score of each candidate category, as well as the category information corresponding to the highest similarity score.
[0112] After the local model outputs the recognition results, it compares the highest similarity score in the results with a preset recognition threshold. If the highest similarity score is greater than or equal to the preset threshold, the target spatial icon is considered successfully recognized; otherwise, if the highest similarity score is less than the preset threshold, the target spatial icon is considered unrecognized. The preset recognition threshold can be dynamically adjusted based on model accuracy and application scenario requirements. For example, it can be set to 80 points to ensure that only recognition results with sufficiently high similarity scores are adopted, avoiding localization errors caused by misjudgments due to low similarity scores.
[0113] When the highest similarity score in the recognition result is greater than or equal to the preset recognition threshold, and the target control icon is successfully recognized, the third priority positioning method is determined to be available. At this time, it can be determined that the element positioning method of the current touch operation is the edge detection positioning method.
[0114] In this embodiment, edge detection positioning is set as the third priority. When the first two positioning methods fail, the edge detection algorithm extracts the outline features of the control icons from the screenshot and combines them with a lightweight local model for visual recognition and matching. This does not rely on any text information or structured attribute data, effectively covering scenarios with no text or attributes, such as pure icon buttons and custom-drawn controls. It fills the positioning blind spots of the first two levels and provides reliable positioning capabilities for pure icon control scenarios with no text or attributes, thus improving the reliability of script recording.
[0115] If the highest similarity score in the recognition result is less than the preset recognition threshold, or if the feature library is empty and has no registered categories, then the target control icon recognition is determined to fail, and the third priority positioning method is determined to be unavailable.
[0116] If the third priority positioning method is deemed unavailable, the target control icon is input into an external image understanding model for recognition, and a natural language name describing the target control icon is output. The external image understanding model is a large-scale visual understanding model deployed on a remote server or in the cloud, possessing general image semantic understanding capabilities.
[0117] The natural language name output by the external image understanding model can be associated with the image features of the target control icon and registered with the local model, updating the local model's feature library. This allows the local model to directly recognize the same or similar control icons in the future without having to call the external image understanding model again, thus achieving adaptive learning of control icons.
[0118] If the target space icon cannot be recognized by the external image understanding model or the model service is unavailable, the current touch operation can be marked as an abnormal operation. In this case, the element positioning step can be skipped, and manual intervention can be waited for or the next touch operation can be entered into the collection process.
[0119] This embodiment uses a fixed-order, three-level priority strategy to ensure that each valid operation uses the highest-priority available positioning strategy. This fully utilizes the accuracy and stability of native Android system data, effectively covers special interface scenarios such as web page views, self-drawn canvases, and pure icon controls. At the same time, the offline inference capability of the local model ensures the real-time performance of the recording process, and the adaptive learning mechanism of the external model continuously enhances the system's recognition capability with use. Overall, it achieves an organic unity of recording efficiency, positioning accuracy, and scene adaptability.
[0120] S105. Based on the touch operation data, perform element positioning operation through the element positioning method to determine the element positioning parameters corresponding to the current touch operation.
[0121] Continue to refer to Figure 2 and Figure 4 After determining the element positioning method for the current touch operation, the element positioning operation is performed to determine the element positioning parameters.
[0122] In one embodiment, when the element positioning method of the current touch operation is XML attribute positioning method, the target user interface node is repositioned in the XML layout file of the current interface according to the operation coordinates in the touch operation data.
[0123] Among them, the XML attribute positioning method is an element positioning method based on the XML layout file generated by the Android system interface automation testing framework. The XML attribute positioning method utilizes the structured attribute data in the Extensible Markup Language (XML) layout file generated by the Android system interface automation testing framework to establish a mapping relationship between operation coordinates and the attribute information of user interface nodes. Since the attribute information originates from the application's underlying code or system framework and corresponds one-to-one with the interface rendering result, it is unaffected by visual factors such as theme switching, resolution differences, and font changes, and has the highest positioning accuracy and long-term stability.
[0124] Specifically, the operation coordinates and XML layout file are obtained from the touch operation data. Using the operation coordinates as spatial reference points, a hierarchical traversal search is performed in the document object model tree of the XML layout file. The traversal strategy can adopt a top-down depth optimization approach, starting from the root node of the document object model tree and checking the boundary coordinate attributes of each user interface node layer by layer. The boundary coordinate attribute is defined in the form of a rectangular area, containing four values: the x-coordinate of the top left corner, the y-coordinate of the top left corner, the x-coordinate of the bottom right corner, and the y-coordinate of the bottom right corner.
[0125] For each user interface node, determine whether the x-coordinate of the operation coordinate is greater than or equal to the x-coordinate of the top-left corner and less than or equal to the x-coordinate of the bottom-right corner of the node's boundary coordinates, and whether the y-coordinate of the operation coordinate is greater than or equal to the y-coordinate of the top-left corner and less than or equal to the y-coordinate of the bottom-right corner. If the operation coordinate satisfies both of these conditions, the node is included in the candidate node set.
[0126] When the candidate node set contains multiple nodes, the recording program performs candidate node filtering. The filtering rule is to prioritize the leaf node with the smallest area, that is, the deepest node in the Document Object Model tree that no longer contains child nodes, as the target user interface node. If there are still multiple leaf nodes with the smallest area, the node's hierarchy depth in the Document Object Model tree is further compared, and the node with the deepest hierarchy is selected as the target user interface node.
[0127] After locating the target user interface node, extract its attribute set, including identity attributes, text attributes, content description attributes, node type attributes, boundary coordinate attributes, click status attributes, and visibility status attributes. Select the corresponding attribute value from the attribute set as location identification information: if the identity attribute is selected, extract the identity attribute value; if the text attribute is selected, extract the text attribute value; if the content description attribute is selected, extract the content description attribute value. Simultaneously, extract the boundary coordinate attribute of the target user interface node as spatial context information, extract the node type attribute as type context information, and extract the hierarchical path sequence from the root node to the target node as structural context information. Encapsulate the location identification information, spatial context information, type context information, and structural context information into element location parameters. These element location parameters fully describe the semantic identifier, screen position, element type, and document structure position of the target element.
[0128] When the element positioning method of the current touch operation is OCR text positioning, the text content and text position in the current screen screenshot are extracted by optical character recognition based on the current screen screenshot in the touch operation data. Then, by matching the operation coordinates with each text position, the target text content corresponding to the operation coordinates is identified, thereby realizing the positioning of the current touch operation object.
[0129] OCR text localization, in particular, uses optical character recognition technology to extract visual text from a current screenshot and locate elements. It identifies text content from pixel-level image data after interface rendering, matches the identified text and its location information with the user's operation coordinates, and thus determines the target element. This method is suitable for interface types that standard user interface trees cannot cover, such as web views and custom canvas drawing, as well as complex scenarios such as dynamically generated content and cross-platform framework rendering.
[0130] Specifically, when performing element location using OCR text localization, the current screen screenshot and operation coordinates are first obtained from the touch operation data. The optical character recognition engine is then invoked to perform text detection and recognition processing on the current screen screenshot, locating the text blocks within the screenshot and outputting the bounding rectangle coordinates of each text block. Next, the character sequence within each text block is decoded, outputting the recognized text content. Finally, the optical character recognition engine outputs all text content in the current screen screenshot along with its corresponding position information. The position information is represented as a rectangular area, including the x-coordinate of the top-left corner, the y-coordinate of the top-left corner, the x-coordinate of the bottom-right corner, and the y-coordinate of the bottom-right corner.
[0131] The operation coordinates are matched with the location information corresponding to each identified text content. If the operation coordinates are located within the rectangular area of a certain text content, then that text content is identified as the target text content, and the target text content string is encapsulated as location identification information. At the same time, the location information of the target text block, the rectangular area coordinates, is recorded as spatial context information. The location identification information and spatial context information are encapsulated into element location parameters, which describe the content identifier and screen position of the target text content.
[0132] When the element localization method of the current touch operation is edge detection localization, the visual features of the target control icon corresponding to the operation coordinates are extracted from the current screen screenshot based on the current screen screenshot in the touch operation data through the edge detection algorithm. The recognized icon visual features are then matched with the features of known icons through the local model, thereby realizing the recognition of the target control icon.
[0133] Among them, the edge detection localization method is an element localization method based on computer vision edge detection algorithms and a lightweight local image classification model. This localization method extracts the edge contour features of control icons from screenshots and combines them with the visual recognition capabilities of the local model to establish a mapping relationship between user operation coordinates and the recognized icon controls. It is suitable for pure icon control scenarios with no XML attribute information, no text content, or unrecognizable text.
[0134] Specifically, when performing element localization using edge detection, the current screen screenshot and operation coordinates are obtained from the touch operation data. An edge detection algorithm is then used to recognize control icons in the current screenshot, extracting regions with regular geometric shapes as candidate control icons. Feature vectors for each candidate control icon are calculated, such as outline area and aspect ratio. These feature vectors are then matched against a preset control icon feature template to filter out high-confidence control icon regions. Finally, the edge detection algorithm outputs a list of recognized control icons in the current screenshot, along with the location information and edge feature description for each control icon.
[0135] The operation coordinates are matched with the location information corresponding to each control icon. If the operation coordinates fall within the area covered by the location information of a certain control icon, that control icon is identified as the target control icon. Then, the target control icon is cropped and extracted from the current screenshot to form an independent target control icon image, which is then input into the local model for recognition. The local model's feature extraction network extracts high-dimensional abstract feature vectors from the input target control icon image. A classifier network maps these feature vectors to the registered control icon category space, outputting the feature similarity between the target control icon and each registered control icon. The registered control icon with the highest feature similarity is the target control icon.
[0136] Retrieve the identifier name of the registered control icon, encapsulate it into the positioning identifier information of the target control icon, and record the rectangular coordinates of the target control icon's position as spatial context information. Encapsulate the positioning identifier information and spatial context information into element positioning parameters, describing the semantic identifier and screen position of the target control icon.
[0137] It should be noted that the process of determining element positioning parameters is closely linked to the process of determining element positioning methods, but they are logically separate steps. Determining the element positioning method addresses which strategy to use to locate the element, while determining the element positioning parameters addresses which specific identification information to extract. This separation design allows for independent optimization of the selection of positioning strategies and the extraction of positioning information, improving the system's flexibility and maintainability.
[0138] This embodiment executes a customized parameter extraction process for different element location methods, ensuring that each element location method can generate a complete parameter set that matches its characteristics, providing accurate, stable, and reproducible element location basis for subsequent generation of structured natural language operation instructions and script playback.
[0139] S106. Based on the element positioning parameters, generate the structured natural language operation instructions corresponding to the current touch operation to complete the mobile application automated script recording.
[0140] Mobile application automated script recording refers to the process of automatically recording user actions on a mobile application using technical means and converting them into repeatable test scripts. The essential goal of mobile application automated script recording is to transform user actions into a set of machine-understandable and executable instructions. This application achieves mobile application automated script recording by converting the element positioning parameters extracted in the aforementioned steps into structured natural language operation instructions.
[0141] Structured natural language operation instructions are standardized data records composed of a fixed set of fields, simultaneously satisfying human readability and machine parsing requirements. Their natural language aspect lies in the use of everyday human language in the operation description fields to describe the operation behavior, allowing non-technical personnel to directly understand the script's meaning. Their structured nature is reflected in the clear data type and semantic definitions of each field, enabling the playback engine to accurately parse and execute them. This eliminates the technical barriers of code or low-level identifiers in traditional test scripts, allowing non-technical roles such as product managers and business testers to directly participate in script maintenance and optimization.
[0142] In this application, the user performs natural touch operations directly on the mobile terminal. The computer listens to and captures the user's current touch operations in real time through the Android debugging bridge channel, and extracts element positioning parameters through the aforementioned steps. Finally, based on the extracted element positioning parameters, structured natural language operation instructions are generated, thereby realizing automated script recording of mobile applications without requiring the user to have programming skills, without requiring mapping control on the computer, and with a recording experience completely consistent with daily mobile phone use.
[0143] Specifically, all key data associated with the current touch operation is converted into structured natural language instructions, which may include the following fields: operation number, operation type (click / long press / swipe / input), operation target (natural language description), positioning method (XML attribute / OCR text / local model), element positioning parameters, operation coordinates, timestamp, current screenshot path, etc.
[0144] The operation sequence number is the unique serial number corresponding to the current touch operation. This sequence number is generated according to the script recording order, incrementing by one: the first valid touch operation has an operation sequence number of one, and each subsequent valid touch operation's operation sequence number increments by one from the previous one. The operation sequence number identifies the execution order of operations within the overall script and serves as the basic index for sequential execution during playback and for locating operations during script editing.
[0145] The operation type is determined by the recognition results in the aforementioned steps and is limited to three types: click, long press, and swipe. These three operation types cover all touch interaction methods of the mobile application targeted by this application.
[0146] The generation of a natural language description for the operation target relies on the location identifier information in the element location parameter set. Specifically, the corresponding description template can be selected based on the element location method identifier. For example, for XML attribute location, if the location identifier information is a text attribute value, then that text attribute value is directly used as the operation target description, such as "set" or "buy now"; if the location identifier information is an identity identifier attribute value, then the corresponding description is looked up through a preset identity identifier-to-natural language mapping table. If there is no corresponding entry in the mapping table, the identity identifier attribute value itself is used as the description. For OCR text location, the target text content in the location identifier information is used as the operation target description. For edge detection location, the identifier name in the location identifier information is directly used as the operation target description, such as "homepage icon" or "search icon".
[0147] The location method field records a standardized string that identifies the element's location method, including XML attribute location method, OCR text location method, and edge detection location method.
[0148] The location parameter field records the location identifier information in the element's location parameter set, stored in structured data format. For XML attribute location, the location parameter includes the attribute type and attribute value; for OCR text location, the location parameter is the text content string; for edge detection location, the location parameter is the identifier name string.
[0149] The coordinate field records the x and y coordinates of the operation. The timestamp includes the start and end times of the current touch operation. The current screenshot path can include the filename of the current screenshot, which typically contains the operation number and the timestamp information.
[0150] The above fields are assembled into a single structured natural language operation instruction and written into the operation sequence list in the memory buffer. When the user actively ends the recording or the preset automatic saving conditions are met, the recording program writes all the natural language operation instructions in the memory buffer to the persistent storage file.
[0151] In one embodiment, all structured natural language operation instructions corresponding to touch operations are persistently stored in a structured document format. Taking the .xlsx (spreadsheet) format as an example, the advantages of the .xlsx format are that non-technical personnel can directly open and edit it using common office software such as Excel / WPS, it is easy to batch modify the natural language descriptions of operation instructions, it supports version comparison and difference analysis, and it is convenient for team collaboration and script sharing. It is a preferred format that is user-friendly for non-technical personnel.
[0152] For example, when persistently storing structured natural language manipulation instructions using an xlsx file, the worksheet structure is designed as follows: Sheet1 (Operation Sequence): Contains columns such as sequence number, operation type, operation description (natural language), positioning method, positioning parameters, X coordinate, Y coordinate, waiting time, and screenshot file name.
[0153] Sheet2 (metadata): Records basic script information such as application package name, application version, device model, system version, screen resolution, recording time, and the person who recorded it.
[0154] Sheet3 (Screenshot Index): Records the screenshot file path and image fingerprint hash value corresponding to each operation.
[0155] Understandably, the structured document format used for persistent storage of structured natural language operation instructions is not limited to xlsx (spreadsheet) format. Other structured document formats can also be used, such as: xlsx (spreadsheet), JSON (JavaScript Object Notation), YAML (AML Ain't Markup Language), CSV (Comma-Separated Values), XML (Extensible Markup Language), etc. The specific structured storage format used can be determined based on the actual application scenario and requirements.
[0156] After all structured natural language operation instructions have been persistently stored, the current mobile application automation script recording process ends.
[0157] This embodiment eliminates the technical barriers of traditional test scripts by converting element positioning parameters into structured natural language operation instructions, enabling non-technical personnel to directly understand, edit, and maintain script content, significantly reducing script maintenance costs. Persistent storage of the structured natural language operation instructions using a structured storage format ensures script integrity and traceability, ultimately achieving a complete closed loop from natural user operation to reusable, collaborative, and maintainable automated test scripts.
[0158] This embodiment provides a method for automated script recording in mobile applications. This method establishes an Android debugging bridge communication connection and monitors the system's underlying touch event stream in real time, directly capturing user touch operation events on the mobile terminal. This avoids behavioral deviations caused by mapping operations on the computer, ensuring the authenticity of the recorded data. Touch operation events are identified based on preset time intervals and coordinate offset judgment conditions, ensuring the accuracy of operation type determination. While identifying the operation type, multi-dimensional touch operation data is extracted, providing a complete digital description foundation for subsequent processing. By comparing screenshots before and after the operation, the validity of the current touch operation is determined, automatically filtering out invalid operations such as accidental touches and blank clicks, avoiding redundant steps in the script, and improving script recording efficiency and data validity. After confirming the operation's validity, the element positioning method is determined based on a preset multi-level element positioning priority strategy, ensuring that each operation uses the highest priority available positioning strategy, improving the efficiency and stability of element positioning operations. The system performs element positioning operations based on the element positioning method, extracts element positioning parameters, and converts the current touch operation into structured natural language operation instructions based on the element positioning parameters to complete script recording. This allows non-technical personnel to directly understand and maintain the script content, while ensuring the reproducibility of the current touch operation and improving the accuracy of automated script recording results.
[0159] Please see Figure 6 , Figure 6 This is a schematic diagram of the structure of a first embodiment of a mobile application automation script recording device provided in this application. The mobile application automation script recording device is used to execute the aforementioned mobile application automation script recording method.
[0160] like Figure 6 As shown, the mobile application automated script recording device 200 includes: an event listening module 201, a touch operation recognition module 202, a touch operation validity judgment module 203, an element positioning method determination module 204, an element positioning module 205, and a structured data generation module 206.
[0161] Event listening module 201 is used to listen for touch operation events triggered by the user on the mobile terminal through ADB log listening service after the computer and the mobile terminal establish an ADB communication connection. The touch operation recognition module 202 is used to identify the current touch operation corresponding to the touch operation event based on at least one preset touch operation judgment condition, determine the operation type of the current touch operation, and extract the touch operation data corresponding to the current touch operation. The current touch operation data includes the timestamp of the current touch operation, operation coordinates, the XML layout file of the current interface of the mobile terminal, and the current screen screenshot. The touch operation validity determination module 203 is used to compare the current screen screenshot with the previous screen screenshot before the current touch operation occurred to determine the validity of the current touch operation; The element positioning method determination module 204 is used to determine the element positioning method of the current touch operation based on a preset multi-level element positioning priority strategy when the current touch operation is determined to be a valid operation. Element positioning module 205 is used to perform element positioning operation based on the touch operation data and through the element positioning method to determine the element positioning parameters corresponding to the current touch operation; The structured data generation module 206 is used to generate structured natural language operation instructions corresponding to the current touch operation based on the element positioning parameters, and to complete the automated script recording of the mobile application.
[0162] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the device and each module described above can be referred to the corresponding processes in the aforementioned embodiments of the mobile application automated script recording method, and will not be repeated here.
[0163] The apparatus provided in the above embodiments can be implemented as a computer program, which can be used in, for example... Figure 7 It runs on the computer device shown.
[0164] Please see Figure 7 , Figure 7 This is a schematic block diagram illustrating the structure of a computer device according to an embodiment of this application. The computer device may be a server.
[0165] See Figure 7 The computer device includes a processor, memory, and network interface connected via a system bus, wherein the memory may include non-volatile storage media and internal memory.
[0166] Non-volatile storage media can store operating systems and computer programs. These computer programs include program instructions that, when executed, cause the processor to perform any method of recording automated scripts for mobile applications.
[0167] The processor provides computing and control capabilities, supporting the operation of the entire computer device.
[0168] Internal memory provides an environment for the execution of computer programs stored on non-volatile storage media. When executed by a processor, the computer program enables the processor to execute any method of recording automated scripts for mobile applications.
[0169] This network interface is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that... Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0170] It should be understood that the processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.
[0171] In one embodiment, the processor is configured to run a computer program stored in memory to perform the following steps: After the computer establishes an ADB communication connection with the mobile terminal, it listens for touch operation events triggered by the user on the mobile terminal through the ADB log listening service. Based on at least one preset touch operation determination condition, the current touch operation corresponding to the touch operation event is identified, the operation type of the current touch operation is determined, and the touch operation data corresponding to the current touch operation is extracted. The current touch operation data includes the timestamp of the current touch operation, operation coordinates, the XML layout file of the current interface of the mobile terminal, and the current screen screenshot. The validity of the current touch operation is determined by comparing the current screenshot with the previous screenshot taken before the current touch operation occurred. When the current touch operation is determined to be a valid operation, the element positioning method of the current touch operation is determined based on a preset multi-level element positioning priority strategy; Based on the touch operation data, an element positioning operation is performed using the element positioning method to determine the element positioning parameters corresponding to the current touch operation; Based on the element positioning parameters, a structured natural language operation instruction corresponding to the current touch operation is generated to complete the automated script recording for the mobile application.
[0172] The embodiments of this application also provide a computer-readable storage medium storing a computer program, the computer program including program instructions, and the processor executing the program instructions to implement any of the mobile application automated script recording methods provided in the embodiments of this application.
[0173] The computer-readable storage medium can be an internal storage unit of the computer device described in the foregoing embodiments, such as the hard disk or memory of the computer device. The computer-readable storage medium can also be an external storage device of the computer device, such as a plug-in hard disk, SmartMediaCard (SMC), SecureDigital (SD) card, or FlashCard equipped on the computer device.
[0174] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for recording automated scripts for mobile applications, characterized in that, Applied to a computer, the method includes: After the computer establishes an ADB communication connection with the mobile terminal, it listens for touch operation events triggered by the user on the mobile terminal through the ADB log listening service. Based on at least one preset touch operation determination condition, the current touch operation corresponding to the touch operation event is identified, the operation type of the current touch operation is determined, and the touch operation data corresponding to the current touch operation is extracted. The current touch operation data includes the timestamp of the current touch operation, operation coordinates, the XML layout file of the current interface of the mobile terminal, and the current screen screenshot. The validity of the current touch operation is determined by comparing the current screenshot with the previous screenshot taken before the current touch operation occurred. When the current touch operation is determined to be a valid operation, the element positioning method of the current touch operation is determined based on a preset multi-level element positioning priority strategy; Based on the touch operation data, an element positioning operation is performed using the element positioning method to determine the element positioning parameters corresponding to the current touch operation; Based on the element positioning parameters, a structured natural language operation instruction corresponding to the current touch operation is generated to complete the automated script recording for the mobile application.
2. The mobile application automated script recording method according to claim 1, characterized in that, The touch operation events include consecutive and adjacent press events and release events; The method of listening to touch operation events triggered by the user on the mobile terminal through the ADB log listening service includes: The ADB log listening service is used to listen to the system-level touch event stream of the mobile terminal. When a touch event that touches the screen is detected in the underlying touch event stream of the system, the start time and starting coordinates of the current touch operation are marked. After the press event, the touch point coordinate changes are continuously tracked until a lift event is detected that removes the touch from the screen, at which point the termination time and end coordinates of the current touch operation are marked.
3. The mobile application automated script recording method according to claim 1, characterized in that, The step of identifying the current touch operation corresponding to the touch operation event based on at least one preset touch operation determination condition, and determining the operation type of the current touch operation, includes: Obtain the start time, end time, start coordinates, and end coordinates of the current touch operation; Calculate the time interval between the start time and the end time, and calculate the coordinate offset of the start coordinate and the end coordinate; When the time interval and the coordinate offset simultaneously meet the determination conditions of any preset touch operation, the operation type of the current touch operation is determined.
4. The mobile application automated script recording method according to claim 3, characterized in that, When the time interval and the coordinate offset simultaneously satisfy the determination condition of any preset touch operation, the operation type of the current touch operation is determined, including: When the time interval is within the first interval range and the coordinate offset is within the first offset distance range, the operation type of the current touch operation is determined to be a click operation; When the time interval is within the second interval range and the coordinate offset is within the second offset distance range, the operation type of the current touch operation is determined to be a long press operation; When the time interval is within the third interval range and the coordinate offset is within the third offset distance range, the operation type of the current touch operation is determined to be a swipe operation.
5. The mobile application automated script recording method according to claim 1, characterized in that, The method for determining the element positioning of the current touch operation based on a preset multi-level element positioning priority strategy includes: The positioning methods are determined sequentially according to the first priority positioning method, the second priority positioning method, and the third priority positioning method. The first positioning method that is determined to be usable is determined as the element positioning method for the current touch operation.
6. The mobile application automated script recording method according to claim 5, characterized in that, The first priority location method is the XML attribute location method; The method for determining the element positioning of the current touch operation based on a preset multi-level element positioning priority strategy includes: Based on the operation coordinates in the touch operation data, search for the corresponding target user interface node in the XML layout file of the current interface; Detect whether the target user interface node contains valid attribute information, wherein the attribute information includes at least one of identity attribute, text attribute, and content description attribute; When the target user interface node contains valid attribute information, it is determined that the first priority positioning method is available, and the element positioning method of the current touch operation is determined to be the XML attribute positioning method.
7. The mobile application automated script recording method according to claim 5, characterized in that, The second priority positioning method is OCR text positioning. The method for determining the element positioning mode of the current touch operation based on a preset multi-level element positioning priority strategy further includes: When the first priority positioning method is unavailable, OCR text recognition processing is performed on the current screenshot to obtain the text content in the current screenshot and the location information corresponding to the text content. Based on the operation coordinates of the current touch operation and the position information, determine whether there is corresponding target text content at the operation coordinate position; When the target text content exists at the operation coordinate position, it is determined that the second priority positioning method is available, and the element positioning method of the current touch operation is determined to be the OCR text positioning method.
8. The mobile application automated script recording method according to claim 1, characterized in that, The third priority positioning method is the edge detection positioning method; The method for determining the element positioning mode of the current touch operation based on a preset multi-level element positioning priority strategy further includes: When the second priority positioning method is unavailable, edge detection is performed on the current screenshot to identify the control icons in the current screenshot and the position information corresponding to the control icons; The target control icon is determined based on the operation coordinates of the current touch operation and the position information of the control icon; The target control icon is input into the local model for recognition, and the recognition result of the local model is obtained; When the recognition result indicates that the target control icon has been successfully recognized, it is determined that the third priority positioning method is available, and the element positioning method of the current touch operation is determined to be the edge detection positioning method.
9. A mobile application automated script recording device, characterized in that, The mobile application automated script recording device includes: The event listening module is used to listen for touch operation events triggered by the user on the mobile terminal through the ADB log listening service after the computer and the mobile terminal establish an ADB communication connection. The touch operation recognition module is used to identify the current touch operation corresponding to the touch operation event based on at least one preset touch operation judgment condition, determine the operation type of the current touch operation, and extract the touch operation data corresponding to the current touch operation. The current touch operation data includes the timestamp of the current touch operation, operation coordinates, the XML layout file of the current interface of the mobile terminal, and the current screen screenshot. The touch operation validity determination module is used to compare the current screenshot with the previous screenshot before the current touch operation occurred to determine the validity of the current touch operation; The element positioning method determination module is used to determine the element positioning method of the current touch operation based on a preset multi-level element positioning priority strategy when the current touch operation is determined to be a valid operation. The element positioning module is used to perform an element positioning operation based on the touch operation data and through the element positioning method to determine the element positioning parameters corresponding to the current touch operation. The structured data generation module is used to generate structured natural language operation instructions corresponding to the current touch operation based on the element positioning parameters, and to complete the automated script recording of the mobile application.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, it implements the steps of the mobile application automated script recording method as described in any one of claims 1 to 8.