Vision-based device testing method and apparatus, storage medium, and electronic device
By correcting the distortion of the test frame image sequence, a corrected image sequence is generated, which solves the problem of test accuracy caused by camera displacement or screen tilt, and achieves more accurate equipment test results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI ANQINZHIXING AUTOMOTIVE ELECTRONICS CO LTD
- Filing Date
- 2026-03-27
- Publication Date
- 2026-07-03
Smart Images

Figure CN122335706A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of vision technology, and more specifically, to a vision-based device testing method and apparatus, storage medium, and electronic device. Background Technology
[0002] In the field of visual performance testing, for cockpit function testing and interface response smoothness performance testing, such as the startup response time and interface switching speed of applications (APPs) on the display screen, the test scheme of "fixed camera shooting the cockpit display screen + preset region of interest parameters" is widely used. However, the above-mentioned related technologies have the problem of low test accuracy. Summary of the Invention
[0003] This application provides a vision-based device testing method and apparatus, storage medium and electronic device, to at least solve the technical problem of low testing accuracy in related technologies.
[0004] According to one aspect of the embodiments of this application, a vision-based device testing method is provided, comprising: during the testing of a device under test, acquiring a sequence of test frame images corresponding to a test interface displayed on the screen of the device under test, wherein the test frame image sequence is obtained by continuously acquiring images of the test interface through an image acquisition device; extracting the display area where the display screen is located in each test frame image in the test frame image sequence, and performing distortion correction on the display area in each test frame image to obtain a corrected image sequence, wherein the corrected image sequence includes a corrected image corresponding to each test frame image; identifying a start frame image and an end frame image corresponding to a test scene in the corrected image sequence, and determining the test result of the test scene based on the timestamp information of the start frame image and the timestamp information of the end frame image.
[0005] According to another aspect of the embodiments of this application, a vision-based device testing apparatus is also provided, comprising: an acquisition unit, configured to acquire a sequence of test frame images corresponding to a test interface displayed on the display screen of the device under test during the testing process, wherein the test frame image sequence is obtained by continuously acquiring images of the test interface through an image acquisition device; a correction unit, configured to extract the display screen area where the display screen is located in each test frame image in the test frame image sequence, and perform distortion correction on the display screen area in each test frame image to obtain a corrected image sequence, wherein the corrected image sequence includes a corrected image corresponding to each test frame image; and a determination unit, configured to identify a start frame image and an end frame image corresponding to a test scene in the corrected image sequence, and determine the test result of the test scene based on the timestamp information of the start frame image and the timestamp information of the end frame image.
[0006] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, wherein a computer program is stored therein, wherein the computer program is configured to perform the steps in any of the above method embodiments when executed by a processor.
[0007] According to another aspect of the embodiments of this application, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to perform the steps in any of the method embodiments described above.
[0008] According to another aspect of the embodiments of this application, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to perform the steps of any of the above method embodiments through the computer program.
[0009] This application enables the acquisition of a sequence of test frame images corresponding to the tested interface on the display screen of the device under test during the testing process. The display screen area in each test frame image is extracted and distortion corrected to obtain a corrected image sequence. This eliminates image distortion caused by image acquisition device position offset, display screen tilt, or optical distortion, ensuring that the geometric structure of the display screen area remains consistent in each frame image. Consequently, the recognition process of the starting and ending frame images is based on stable, distortion-free visual features, avoiding area positioning offset or feature misjudgment caused by image deformation. Finally, the test results are directly calculated based on the timestamps of the starting and ending frame images, ensuring that the accuracy of the test results is unaffected by external acquisition environment disturbances. This significantly improves the accuracy of vision-based recognition tests, thus solving the technical problem of low test accuracy in related technologies. Attached Figure Description
[0010] Figure 1 This is a schematic diagram illustrating an application scenario of a vision-based device testing method according to an embodiment of this application;
[0011] Figure 2 This is a flowchart illustrating an optional vision-based device testing method according to an embodiment of this application;
[0012] Figure 3 This is a flowchart of an optional method for determining the start frame and the end frame according to an embodiment of this application;
[0013] Figure 4 This is a flowchart of an optional corner recognition method according to an embodiment of this application;
[0014] Figure 5 This is a flowchart of an optional perspective transformation according to an embodiment of this application;
[0015] Figure 6 This is an optional test flowchart according to an embodiment of this application;
[0016] Figure 7 This is a structural block diagram of an optional vision-based device testing apparatus according to an embodiment of this application;
[0017] Figure 8 This is a computer system architecture block diagram of an optional electronic device according to an embodiment of this application. Detailed Implementation
[0018] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0019] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0020] According to one aspect of the embodiments of this application, a vision-based device testing method is provided. Optionally, in this embodiment, the above-described vision-based device testing method may be applied, but is not limited to, to applications such as... Figure 1 The hardware environment shown includes an image acquisition device 102, a device under test (DUT) 104, and an electronic device 106. The electronic device 106 can connect to the image acquisition device 102 via a network. The image acquisition device 102 captures the DUT interface displayed on the screen of the DUT 104 and uploads the acquired data to the electronic device 106. The electronic device 106 performs vision-based device testing on the DUT 104 based on the data returned by the image acquisition device 102.
[0021] The aforementioned networks may include, but are not limited to, at least one of the following: wired networks and wireless networks. The aforementioned wired networks may include, but are not limited to, at least one of the following: wide area networks (WANs), metropolitan area networks (MANs), and local area networks (LANs). The aforementioned wireless networks may include, but are not limited to, at least one of the following: Wireless Fidelity (WIFI) and Bluetooth. The image acquisition device 102 may include devices such as cameras and visual sensors, and the electronic device 106 may be, but is not limited to, personal computers (PCs), mobile phones, tablet computers, etc.
[0022] The vision-based device testing method of this application embodiment can be executed by electronic device 106. Taking the execution of the vision-based device testing method of this embodiment by electronic device 106 as an example, Figure 2 This is a flowchart illustrating an optional vision-based device testing method according to an embodiment of this application, such as... Figure 2 As shown, the process of this method may include the following steps S202 to S206.
[0023] Step S202: During the testing of the device under test, a sequence of test frame images corresponding to the interface under test displayed on the screen of the device under test is acquired. The sequence of test frame images is obtained by continuously acquiring images of the interface under test through an image acquisition device.
[0024] Step S204: Extract the display area where the display screen is located in each test frame image in the test frame image sequence, and perform distortion correction on the display area in each test frame image to obtain a corrected image sequence, wherein the corrected image sequence includes the corrected image corresponding to each test frame image.
[0025] Step S206: Identify the start frame image and end frame image corresponding to the scene under test in the corrected image sequence, and determine the test result of the scene under test based on the timestamp information of the start frame image and the timestamp information of the end frame image.
[0026] The vision-based device testing method in this embodiment can be applied to the field of vision and to scenarios where devices are tested based on vision. Optionally, it can be applied to scenarios where the startup response time of smart cockpit apps (Applications) is tested, to automatically measure the response time of navigation, music, and other apps on the in-vehicle central control screen from the moment they are clicked to the moment they are fully loaded; or it can be applied to scenarios where the smoothness of smart cockpit interface switching is tested, to evaluate the frame rate and stuttering of the vehicle system when switching between different functional interfaces (such as air conditioning → radio → Bluetooth); or it can also be applied to scenarios where the screen interaction of smart wearable devices is tested, to measure the visual response speed of smartwatches or bracelets when launching applications and switching menus.
[0027] Taking intelligent cockpit visual performance testing as an example, for cockpit function testing and interface response smoothness performance testing, such as application (APP) startup response time and interface switching speed, a testing scheme of "fixed camera shooting the cockpit display screen + preset Region of Interest (ROI) parameters" is widely adopted. Its core logic is: pre-set the ROI parameters of the target area on the display screen (such as APP icon, loading screen), continuously capture images with the camera, match the ROI parameters to identify the test start frame (such as the initial state of the interface after the APP startup command is triggered) and the end frame (such as the interface after the APP is fully loaded), and then calculate the performance time. Here, the target area refers to the pre-defined region of interest; for example, the target area may include the APP icon, loading screen, etc.
[0028] However, the aforementioned technologies have the following common defects in practical engineering applications: lack of dynamic calibration and deformation correction mechanisms, poor anti-interference ability of ROI parameters, that is, the ROI parameters of the existing solutions are preset at one time. Once there are slight changes such as slight camera displacement (±5mm or more), screen tilt (±3° or more), or image distortion (trapezoidal, parallelogram) during the test, the preset ROI parameters cannot match the actual target area of the image, resulting in failure of the recognition of the start frame image / end frame image, interruption or deviation of performance timing; further affecting the accuracy of various test results.
[0029] To at least partially address the aforementioned technical problem of low test accuracy due to poor parameter anti-interference, this embodiment acquires a sequence of test frame images corresponding to the tested interface on the device under test's display screen during the testing process. The display screen area in each test frame image is extracted and distortion corrected to obtain a corrected image sequence. This eliminates image distortion caused by image acquisition device position offset, display screen tilt, or optical distortion, ensuring that the geometric structure of the display screen area remains consistent in each frame. This allows the recognition process of the starting and ending frame images to be based on stable, distortion-free visual features, avoiding... This method avoids regional positioning offsets or feature misjudgments caused by image distortion. The test results are directly calculated based on the timestamps of the starting and ending frames, ensuring that the accuracy of the test results is not affected by external acquisition environment disturbances. This significantly improves the accuracy of vision-based recognition tests and solves the problem in related technologies where display area extraction relies solely on simple boundary recognition without distortion correction steps, leading to regional positioning deviations and affecting test accuracy. Furthermore, it can solve the problem that preset ROI parameters become invalid and cause test interruptions when the image acquisition device (such as a camera) is slightly offset, the display screen is tilted, or the captured image produces trapezoidal / parallelogram distortion.
[0030] In this embodiment, the device under test is the device that needs to be tested. For example, the device under test can be a smart cockpit, a smartphone, or a smart wearable device. The device under test can also be other devices with a display screen, and no specific limitations are made here.
[0031] The tested scenario refers to a user interface interaction behavior or functional response process with a start state and an end state that needs to be quantified and evaluated visually. Optionally, the time consumed from the start state to the end state of the tested scenario can be measured objectively, accurately, and automatically through external visual acquisition methods, i.e., visual response latency. For example, the tested scenario can be a scenario that tests the interface switching speed or a scenario that tests the pop-up response time.
[0032] The interface under test refers to the interface on the display screen of the device under test that is being monitored and evaluated. Optionally, the interface under test can correspond to the display screen of a specific functional scenario, such as the APP startup screen (such as the black screen / loading icon screen when the navigation application is first clicked to start) to the main screen after the APP is fully loaded; the pop-up screen of the air conditioning temperature adjustment panel in the vehicle system; the switching screen of the music player; the transition screen of the central control screen from the main menu to the settings menu; the screen of the appearance and disappearance of system pop-ups (such as "Bluetooth connection successful").
[0033] The test frame image sequence is an image sequence obtained by continuously acquiring images of the interface under test through an image acquisition device. The image acquisition device refers to a visual device used to continuously acquire the screen image of the device under test. For example, the image acquisition device can be a camera or a visual sensor.
[0034] Optionally, images of the display area of the device under test can be captured at a fixed frequency using a camera or vision sensor to form a series of images arranged in chronological order, i.e., a test frame image sequence. This test frame image sequence can completely record the visual evolution process of the interface displayed on the screen of the device under test from the initial state to the target state.
[0035] The test frame image sequence includes multiple test frame images. Each test frame image in the test frame image sequence refers to an original image frame containing the complete screen of the device under test, continuously acquired by an image acquisition device at a fixed sampling frequency during the test of the device under test (DUT) in the test scenario. In this embodiment, the screen of the DUT in each test frame image in the test frame image sequence acquired by the image acquisition device may have geometric distortion problems caused by shooting angle deviation. For example, when there is a non-orthogonal installation relationship between the image acquisition device (such as a camera) and the screen of the DUT, the screen may appear as a trapezoid, parallelogram, or irregular quadrilateral in each test frame image, rather than a true rectangle. Small displacements of the image acquisition device (±5mm or more) or installation tilt of the screen (±3° or more) will further aggravate the perspective distortion of the screen boundary in the image. Therefore, it is necessary to correct each test frame image in the test frame image sequence before testing.
[0036] The display area in each test frame image refers to the area enclosed by the physical boundaries of the display screen in each test frame image. Optionally, the display area in each test frame image acquired by the image acquisition device can be rectangular, parallelogram-shaped, or curved.
[0037] In some embodiments, edge detection operators can be used to perform Canny edge extraction on each test frame image in the test frame image sequence. Four main edge lines are detected using Hough transform, the coordinates of quadrilateral vertices are fitted, and the display contours that meet the constraints of the angles between vertices and the ratio of side lengths are selected. The quadrilateral region enclosed by these contours is then output as the display area. Alternatively, a pre-trained semantic segmentation neural network model can be used to perform pixel-level classification on each test frame image in the test frame image sequence, identify and output a mask image of the display area, and obtain the complete boundary of the display area through mask dilation and contour closure processing. This yields the display area where the display is located for each test frame image in the test frame image sequence.
[0038] A corrected image sequence refers to a new image sequence generated after distortion correction processing, corresponding one-to-one with the original test frame image sequence. The corrected image sequence includes the corrected image corresponding to each test frame image. Distortion correction refers to the process of restoring an ideal projection (such as trapezoidal distortion) caused by the shooting angle or tilt of the image acquisition device to a normal viewing plane image through image geometric transformation technology, thus eliminating the influence of geometric distortion. In each corrected image in the corrected image sequence, the display area is restored to a standard rectangular shape.
[0039] In some embodiments, the distortion correction of the display area in each test frame image to obtain a corrected image sequence may include the following process: extracting the quadrilateral boundary of the display screen using Canny edge detection and Hough transform, calculating the perspective transformation matrix based on the coordinates of the four corner points to obtain a standard rectangular coordinate system, mapping the display area in each test frame image to the standard rectangular coordinate system to obtain multiple corrected image frames, and assembling the corrected image sequence in chronological order.
[0040] In some embodiments, the distortion correction of the display area in each test frame image to obtain a corrected image sequence may further include the following process: extracting the boundary corner points of the display area of each test frame image to form an irregular quadrilateral; then, according to the physical resolution of the display screen of the device under test, predefining the pixel size and center position of the target rectangle; subsequently, using the Direct Linear Transform (DLT) algorithm to solve the 3×3 perspective transformation matrix from the source corner point to the target rectangle corner point, and using the 3×3 perspective transformation matrix to perform coordinate inverse mapping and pixel resampling on the pixels in the display area of the test frame image, outputting the corrected image; performing the same operation on each test frame image in the test frame image sequence in sequence to obtain a set of corrected image sequences.
[0041] The corrected images in the corrected image sequence are all directly derived from each test frame image in the test frame image sequence. The generation process of the corrected images in the corrected image sequence does not change the timestamp order of the images, thus maintaining the temporal integrity of the test frame image sequence. The visual content (such as interface elements, text, icons, etc.) carried by the corrected images is completely consistent with the corresponding test frame images, and only the spatial geometry has been transformed.
[0042] It should be noted that the test frame images are raw acquired data. Due to the shooting angle deviation, the display area in the test frame images appears as a non-rectangular distortion shape (such as trapezoids or parallelograms), and the pixel coordinates do not have an orthogonal correspondence with the real physical plane, resulting in spatial distortion in area positioning. The corrected images, on the other hand, are geometrically normalized results generated after correcting the test frame images. The display area in the corrected images is reconstructed into a standard rectangle, and the pixel distribution conforms to the projection relationship under an ideal frontal viewing angle, eliminating geometric distortion caused by image acquisition device offset or screen tilt. It can be understood that each test frame image in the test frame image sequence includes the display area of the device under test and other component areas of the device under test, while the corrected image corresponding to each test frame image in the corrected image sequence only includes the corrected standard rectangular display area.
[0043] The starting frame image refers to the image frame in the test frame image sequence that is the first visually recognizable change in the interface after the function under test (such as APP startup) is triggered; the ending frame image refers to the image frame in the test frame image sequence that is the image frame in the interface under test that has reached a stable target state when the function under test is completed.
[0044] Optionally, by comparing the pixel change rate of each frame of the corrected image in the corrected image sequence with that of a preset reference image, when the pixel change rate in N consecutive corrected images exceeds a first preset threshold, the first frame of the corrected image in the N consecutive corrected images is determined to be the starting frame image, where N is a positive integer greater than or equal to 2; when the pixel change rate in M consecutive corrected images is continuously lower than a second preset threshold and the target region features are stable, the first frame of the corrected image in the M consecutive corrected images is determined to be the ending frame image. Alternatively, by detecting the average brightness and edge gradient change trend of the target region of each corrected image in the corrected image sequence, when the average brightness of the first corrected image in the corrected image sequence exceeds the third preset threshold and the sum of edge gradients exceeds the fourth preset threshold, the first corrected image is determined to be the starting frame image; when the average brightness of the corrected images in the corrected image sequence for a continuous Q-frame is stable within the first preset range and the sum of edge gradients is less than the fifth preset threshold, the first frame image in the continuous Q-frame corrected images in the corrected image sequence is determined to be the ending frame image, where Q is a positive integer greater than or equal to 2.
[0045] The timestamp information of the starting frame image refers to the time stamp information corresponding to the starting frame image that captures the first identifiable visual change in the target area of the image; the timestamp information of the ending frame image refers to the time stamp information corresponding to the ending frame image that captures the target area after all loading has been completed and the tested interface has reached a stable final state.
[0046] Optionally, determining the test result of the tested scenario based on the timestamp information of the start frame image and the timestamp information of the end frame image may include the following process: In the case that the tested scenario is the application startup scenario, calculate the difference between the timestamp of the end frame image and the timestamp of the start frame image, use the difference in timestamps as the application startup response time, thereby obtaining the test result of the tested scenario, and determining whether the tested scenario of application startup meets the quantitative indicators of performance evaluation.
[0047] For example, record the start and end frame timestamps (that is, record the timestamp information of the start frame image and the end frame image), calculate the difference between the timestamp information of the start frame image and the timestamp information of the end frame image as a performance indicator (such as startup response time), and automatically store it in the database to complete the single-scenario test.
[0048] Through the embodiments provided in this application, during the testing of the device under test, a sequence of test frame images corresponding to the interface under test on the display screen of the device under test is acquired, and the display screen area in each test frame image is extracted and distortion corrected to obtain a corrected image sequence. This can eliminate the image distortion caused by the positional offset of the image acquisition device, the tilt of the display screen, or optical distortion, thereby ensuring that the geometric structure of the display screen area in each frame image remains consistent. This allows the recognition process of the start frame image and the end frame image to be based on stable, distortion-free visual features, avoiding regional positioning offset or feature misjudgment caused by image deformation. Finally, the test result is directly calculated based on the timestamp of the start frame image and the timestamp of the end frame image, so that the accuracy of the test result is not affected by external acquisition environment disturbances, significantly improving the test accuracy based on visual recognition. Therefore, it can solve the technical problem of low test accuracy in related technologies.
[0049] In an exemplary embodiment, the method further includes: acquiring an original frame image corresponding to an initial interface displayed on the display screen of the device under test, wherein the original frame image is obtained by acquiring an image of the initial interface using an image acquisition device; extracting the display screen area where the display screen is located in the original frame image, and performing distortion correction on the display screen area in the original frame image to obtain an image to be verified; mapping the feature region parameters of a target reference image onto the image to be verified to obtain a verification reference image, wherein the target reference image is a reference image with pre-calibrated feature regions corresponding to the scene under test, and the feature region parameters are used to indicate the calibrated feature regions in the target reference image; performing image verification on the verification reference image based on the target reference image to obtain an image verification result of the image to be verified, wherein acquiring the test frame image sequence is performed when the image verification result is a successful verification.
[0050] In related technologies, the lack of a pre-test validation step for parameter validity leads to insufficient reliability of test results. Specifically, existing technologies do not verify the validity of preset parameters before test execution, directly using historical parameters. If these parameters have become invalid due to hardware changes or software interface updates, it can cause errors in identifying key test nodes, ultimately resulting in distorted test data that fails to objectively reflect the product's true state. Furthermore, parameter reusability is low, and cross-scenario adaptation costs are high. Additionally, after hardware installation and fine-tuning or software interface updates to the device under test (e.g., a smart cockpit), old ROI parameters become incompatible with the new scenario, requiring manual recalibration, increasing test preparation time, and reducing the efficiency of cross-model and cross-version testing. To address the aforementioned issues, geometric compensation for differences in shooting perspectives is achieved through display area extraction and distortion correction of the original frame image. By mapping pre-calibrated feature region parameters to the corrected image and performing visual consistency verification, the validity of test parameters in the current environment is automatically verified. This effectively identifies parameter failures caused by image acquisition device displacement, screen tilt, or minor adjustments to the installation position, thus blocking invalid testing processes before test execution and avoiding frame recognition errors and performance data distortion caused by parameter mismatch. Different target reference images are preset for different test scenarios; when the tested scenario is incompatible, a target reference image matching the tested scenario can be quickly replaced, eliminating the need for manual calibration. This solves the problem of missing parameter validity verification before testing, where parameters are still used even after hardware displacement or software interface updates, leading to test node recognition errors and data distortion.
[0051] In this embodiment, the original frame image is an image obtained by capturing the initial interface using an image acquisition device. The initial interface refers to the visual interface that is stably displayed on the screen of the device under test before the test scenario is triggered, representing the system or application in a standby or ready state. Optionally, the initial interface is the static stable state when the function under test is not triggered, while the interface under test is the intermediate or final state that changes dynamically during the function response process (such as the loading page during APP startup or the transition frame during interface switching). Furthermore, the content of the initial interface remains unchanged before the test and is used to establish a spatial benchmark. The content of the interface under test undergoes significant visual changes as the test command is executed, and is the target object of the test timing analysis. The initial interface image is used for parameter verification and distortion correction benchmark generation (i.e., generating the corresponding reference of the target benchmark image), while the interface under test image is used to identify the start frame and the end frame, and then calculate the response delay.
[0052] It should be noted that the initial interface is the preceding state of the interface under test. The starting frame of the scene under test can be determined by the first visual change of the initial interface. The initial interface and the interface under test both come from the same display screen in physical space and have the same display area, resolution and installation position. This allows the image to be verified after distortion correction of the initial interface to be effectively mapped and compared with the target reference image in geometric space.
[0053] For example, after the test starts, the display screen of the device under test (smart cockpit) is driven to enter the initial page, and one original initial frame (i.e., original frame image) is captured by the image acquisition device.
[0054] The display area where the display screen is located in the original frame image refers to the area enclosed by the physical boundary of the display screen of the device under test in the original frame image. Optionally, the display area where the display screen is located in the original frame image may be a rectangular or quadrilateral area, or it may be non-rectangular distortion due to the viewing angle of the image acquisition device.
[0055] For example, by using image acquisition devices (such as cameras) deployed in smart cockpits, images can be captured of the initial interface (such as the system homepage) currently displayed on the device under test (such as the vehicle's central control screen) when the test process starts, resulting in an unprocessed raw digital image (i.e., the raw frame image).
[0056] Optionally, the geometric boundary features of the display screen in the original frame image are identified to obtain the display screen area of the original frame image, wherein the geometric boundary features are used to identify the location area of the display screen.
[0057] Alternatively, identifying the geometric boundary features of the display screen in the original frame image may include: performing color space conversion on the original frame image to obtain the converted original frame image; and identifying Z corner points of the display screen in the converted original frame image based on a preset display screen color range, wherein the geometric boundary features include Z corner points, and Z is a positive integer greater than or equal to 2.
[0058] The display screen area in the original frame image suffers from perspective distortion due to the non-perpendicular and non-aligned installation posture between the image acquisition device and the display screen of the device under test. This distortion manifests as a non-rectangular trapezoid, parallelogram, or irregular quadrilateral structure. This distortion alters the geometric relationship between the pixel coordinates of the display screen boundary and the actual physical plane. Directly using the original area for feature matching or ROI mapping will lead to problems such as feature positioning deviation, region overlap misalignment, and parameter failure due to spatial coordinate offset. Furthermore, slight displacement of the image acquisition device, device installation tilt, and changes in environmental viewing angle can further exacerbate these problems, preventing the accurate projection of feature region parameters (such as the position of the APP icon or the loading progress bar area) of the pre-calibrated target reference image onto the image to be verified, resulting in test misjudgment or failure. Therefore, distortion correction of the display screen area in the original frame image is necessary to reconstruct its standard geometric shape from a frontal viewpoint, ensuring that subsequent verification, mapping, and judgment are based on a unified coordinate system.
[0059] The image to be verified is a standardized image generated after distortion correction of the display screen area in the original frame image.
[0060] It should be noted that the image to be verified originates from the display screen area in the original frame image. The pixel data of the image to be verified is generated by correcting the distortion of the display screen area in the original frame image. No external content or artificially synthesized elements are introduced. The image to be verified and the display screen area in the original frame image maintain strict consistency in timestamp, acquisition source, and visual content (such as interface elements, color distribution, and text layout). They represent the same physical scene in different processing stages. It is understandable that the display screen area in the original frame image may exhibit a real-world distortion shape (such as a trapezoid or parallelogram), and the display screen area in the original frame image may be a non-standard rectangle. The image to be verified can be an image of the display screen area in the corrected original frame image. The image to be verified is a standard rectangle, and its geometry has been corrected for a frontal view.
[0061] In some embodiments, edge detection (such as Canny), contour extraction (such as OpenCV's findContours), or corner detection algorithms (such as Harris and Shi-Tomasi) can be used to extract the coordinates of the four corner points of the display area and construct the source quadrilateral; then, according to the nominal aspect ratio of the display, the target corner point coordinates of the ideal rectangle are set, a 3×3 perspective transformation matrix is calculated, the original area is affinely resampled, and the corrected image to be verified is output.
[0062] For example, correcting the distortion of the display screen area in the original frame image to obtain the image to be verified can include: correcting the distortion of the display screen area in the original frame image through perspective transformation to obtain the image to be verified. As another example, the above extraction of the display screen area from the original frame image and the distortion correction of the display screen area in the original frame image can include the following process: performing hue, saturation, and value (HSV) color segmentation and contour detection on the original frame image to identify the corner points of the display screen in the original frame image and segment the effective area; correcting the distortion through perspective transformation to generate the image to be verified.
[0063] A target reference image refers to a static image that is pre-collected and manually calibrated during the initialization phase of the test system for the scenario under test (such as "APP startup" or "interface switching"), representing the ideal test state. Optionally, this target reference image is a screenshot of the display screen without distortion when viewed from the front. The target reference image can be obtained through the Android Debug Bridge (ADB) and can also be called an ADB reference image. Multiple target reference images corresponding to the scenario under test can be pre-stored, with one target reference image corresponding to each scenario under test.
[0064] Optionally, the target reference image includes features such as boundaries, color distribution, and edge gradients. For example, the target reference image includes the central area of the APP icon, the starting pixel block of the loading animation, and the highlighted area of the button after the interface switch is completed.
[0065] Feature region parameters refer to parameters used to quantitatively describe the location and shape of key visual regions in a target reference image. For example, feature region parameters may include coordinate parameters, color / texture parameters, etc. Optionally, feature region parameters may also be metadata describing the location and shape of key regions in the target reference image in the form of coordinates (x, y, width, height), shape (rectangle / polygon), color threshold, etc.
[0066] A verification reference image is an image generated by transforming the feature region parameters in the target reference image to the coordinate system of the image to be verified through geometric mapping. It can be used to compare visual consistency with the image to be verified.
[0067] It should be noted that the verification reference image is generated by mapping the feature region parameters of the target reference image to the spatial coordinates of the image to be verified. It can be the projection of the feature region parameters of the target reference image onto the image to be verified. The image to be verified serves as the geometric spatial basis of the verification reference image, providing an accurate coordinate reference for the feature region parameter mapping. The target reference image provides the source of the feature region parameters for the verification reference image. Understandably, the target reference image can be an image manually calibrated during the initialization phase, containing the complete display screen interface and calibrated feature regions. The image to be verified is generated from the original frame image acquired in the current test environment after distortion correction; it only contains the corrected display screen content and has no labeled areas. The verification reference image is an overlay image generated by mapping the feature region parameters of the target reference image onto the image to be verified; it includes the corrected display screen content and feature regions.
[0068] Optionally, based on the mapping relationship between the pixel coordinate system of the image to be verified after perspective transformation and the standardized coordinate system of the target reference image, the coordinates of the preset ROI rectangular region in the target reference image can be translated and scaled according to the affine transformation matrix to generate a verification reference image spatially aligned with the image to be verified. This verification reference image retains the position, size, and boundary information of the original feature region. Alternatively, the coordinates of the four corner points and the corresponding color feature vectors of the feature region in the target reference image can be extracted. Combined with the corner point registration results obtained from the image to be verified after distortion correction, the feature region parameters can be resampled and mapped along the pixel dimension using a spatial interpolation algorithm to generate a verification reference image that is spatially perfectly aligned with the image to be verified.
[0069] For example, pre-capture the ADB baseline image (including the APP start / stop interface) and corresponding ROI parameters for each test scenario (where the ROI can be pre-labeled in the ADB baseline image), and store the ADB baseline images according to scenario categories; preset the automated operation sequence; after the test starts, drive the display screen of the device under test (such as a smart cockpit) to enter the initial page, collect 1 original initial frame (original frame image), and simultaneously retrieve the ADB baseline image corresponding to the original frame image (denoted as B image) for subsequent processing.
[0070] In this embodiment, image verification of the calibration reference image based on the target reference image can be used to automatically verify whether the pre-calibrated feature region parameters are still applicable to the actual physical environment of the current test scenario before the test is executed. This reduces the occurrence of test misjudgments caused by hardware displacement, screen tilt, camera offset, or interface layout changes from the source. Feature region parameters (such as the position of the APP launch icon) are manually calibrated under ideal conditions (target reference image). However, during actual testing, due to slight movement of the image acquisition device, APP interface updates (such as larger icons, moved progress bar positions), or display distortion, even after correction, the visual content inside the image to be verified may still have a systematic offset or structural mismatch with the original calibration environment. If the old parameters are used directly for frame determination, serious problems will occur, such as the system misidentifying background textures as a loading progress bar or missing key interface changes due to coordinate offset. The verification reference image is a map of the expected position generated by reprojecting the calibration parameters in the target reference image according to the geometric correction relationship of the image to be verified. Image verification between the target reference image and the verification reference image is essentially a spatial consistency comparison. If they are consistent, it means that the current test scene environment has not been significantly disturbed, the parameters are still valid, and the test can be safely performed. If they are inconsistent, it means that the image acquisition device or screen has been offset, or the interface structure has been changed, the parameters are invalid, and the test must be stopped and an alarm must be issued.
[0071] Image verification results characterize whether the feature region parameters of the target reference image are valid and whether the image to be verified can be used as a comparison and correction benchmark for subsequent testing processes. Image verification results refer to the results obtained after performing image verification on the image to be verified and the verification reference image. Image verification results can be used to determine whether to continue executing subsequent performance testing processes. It can be understood that image verification results can be used to determine whether the above steps of obtaining test frame image sequences are successful. Obtaining test frame image sequences is performed when the image verification result is successful.
[0072] Optionally, if the image verification result is that the verification failed, the current test process is terminated and a parameter failure message is output.
[0073] In some embodiments, after acquiring the verification reference image and the image to be verified, the verification reference image and the image to be verified are respectively preprocessed by grayscale conversion and Gaussian filtering. The scale-invariant feature transform (SIFT) feature point detection algorithm is used to extract the feature points of the verification reference image and the image to be verified, and the RANSAC algorithm is used to remove mismatched feature points to build a stable feature correspondence. Based on the feature correspondence, the homography matrix between the verification reference image and the image to be verified is calculated, and the image to be verified is aligned to the coordinate system of the verification reference image by perspective transformation. After alignment, the structural similarity index (SSIM) and mean square error (MSE) of the verification reference image and the image to be verified are calculated pixel by pixel. When SSIM is greater than a first threshold (e.g., 0.92) and MSE is lower than a second threshold (e.g., 50), the image verification result is determined to be "passed verification". Otherwise, a "parameter failure" alarm is output and the test process is terminated.
[0074] In some embodiments, after obtaining the verification reference image and the image to be verified, the coordinates of the four corner points are extracted from the display screen area in the verification reference image and the image to be verified, respectively. The corner points of the image to be verified are mapped to the corner point space of the verification reference image through affine transformation to achieve geometric alignment. After alignment, the ROI regions of the verification reference image and the image to be verified are divided into 16×16 grid blocks. The local texture entropy and average brightness value in each grid block are calculated to construct a feature vector. The cosine similarity of the two sets of feature vectors is compared. If the similarity of all grid blocks is higher than the third threshold (e.g., 0.88) and the average similarity is not lower than the fourth threshold (e.g., 0.91), the image verification result is determined to be "passed verification". Otherwise, it is determined to be "failed verification", and the parameter recalibration prompt mechanism is triggered.
[0075] In this way, by comparing the newly corrected image with the baseline ADB screenshot, the validity of the ROI parameters is determined through feature similarity verification, thus addressing the core pain point of invalid test results due to parameter mismatch and failure at the source. The above correction approach and the baseline comparison process form an automated linkage, which can adapt to changes in environmental location and parameter update scenarios without manual intervention, completely solving the problems of poor adaptability and low reliability of test results in existing solutions.
[0076] This embodiment achieves geometric compensation for differences in shooting perspective by extracting and correcting the display area of the original frame image; by mapping the pre-calibrated feature area parameters to the corrected image and performing visual consistency verification, it achieves automated verification of the validity of test parameters in the current environment; it can effectively identify parameter failures caused by image acquisition device displacement, screen tilt, or fine adjustment of installation position, thereby blocking invalid test processes before test execution and avoiding frame recognition errors and performance data distortion caused by parameter mismatch.
[0077] In an exemplary embodiment, the pixel coordinates of the image to be verified correspond one-to-one with the pixel coordinates of the target reference image, and the resolution of the image to be verified is the same as that of the target reference image. Image verification is performed on the verification reference image based on the target reference image to obtain the image verification result of the image to be verified, including: if the feature similarity between the verification reference image and the target reference image is greater than or equal to a first similarity threshold, determining the image verification result as verified; if the feature similarity between the verification reference image and the target reference image is less than the first similarity threshold, determining the image verification result as verified.
[0078] In this embodiment, during the image verification process, each pixel in the image to be verified is spatially mapped to a pixel in the target reference image at the same position, so that the pixel coordinates of the image to be verified correspond one-to-one with the pixel coordinates of the target reference image. This one-to-one correspondence is achieved by aligning the coordinates after perspective transformation, which can ensure that the image to be verified and the target reference image are completely aligned in the spatial dimension without any rotation, scaling or translation deviation.
[0079] The image size (i.e. resolution) (width × height, in pixels) of the image to be verified is the same as that of the target reference image. This ensures that the image to be verified and the target reference image are perfectly matched in terms of pixel count and density, avoiding feature matching inaccuracies or calculation errors caused by resolution differences.
[0080] For example, a standardized image (i.e., the image to be verified) with the same resolution as the ADB reference image (denoted as image B) corresponding to the original frame image can be generated, which can be denoted as image A, thus facilitating the establishment of the pixel correspondence between the image to be verified and the target reference image.
[0081] Feature similarity is an image matching metric. Optionally, the content consistency of two images can be calculated by extracting and comparing local features (such as key points, descriptors, histograms, structural similarity SSIM, etc.). The numerical range of feature similarity can be 0 to 1. The higher the value of feature similarity, the closer the content of the two images is.
[0082] The first similarity threshold refers to the minimum similarity standard threshold used to determine whether the target reference image and the verification reference image have visual consistency.
[0083] For example, the image to be verified (Image A) and the target reference image (Image B) are called. Based on the feature similarity standard achieved by the two after correction, the ROI parameters defined in Image B are mapped to Image A according to their relative positions. The fit is verified by comparing the feature vectors. If the standard is met, the parameters are determined to be valid, and Image A is used as the frame judgment benchmark; if the standard is not met, an invalidation is indicated and the test is blocked.
[0084] In some embodiments, feature extraction (such as SIFT, ORB, or SSIM features) is performed on the verification reference image and the target reference image, and the similarity values between the verification reference image and the target reference image in local texture, edge distribution, or overall structure are calculated. When the similarity value is not lower than a preset first similarity threshold (e.g., 0.92), it is determined that the content of the verification reference image and the target reference image is highly consistent, indicating that the visual environment on which the ROI parameters depend has not undergone substantial shift, and the image verification result is "passed". When the calculated feature similarity is lower than the first similarity threshold, it indicates that there is a significant difference in visual structure between the image to be verified and the target reference image (e.g., screen content changes, positional shifts exceeding the tolerance range, or distortion correction failure), and the system determines that the current ROI parameters are no longer applicable, the verification result is "failed", thereby preventing the subsequent testing process from continuing.
[0085] Thus, through the above embodiments, there is no need for manual recalibration of ROI parameters. After hardware fine-tuning or APP interface iteration, the system automatically completes parameter verification and frame adaptation. The preparation time for single-scenario testing is greatly shortened, and the efficiency of cross-model and cross-APP version testing is significantly improved. It effectively saves labor and time costs, has a high degree of automation, and greatly reduces adaptation costs. It can solve the problems of low parameter reusability, the need for manual recalibration after hardware fine-tuning or software iteration, high adaptation costs, and low testing efficiency.
[0086] This embodiment achieves an objective and quantitative judgment of image content consistency by setting the image to be verified and the target reference image to correspond one-to-one in pixel coordinates and have the same resolution, and performing binarization verification based on feature similarity and a first similarity threshold. It can effectively identify reference mismatch problems caused by software interface updates or abnormal display content, thereby accurately determining whether the current test parameters are still applicable to the current screen content without manual intervention, avoiding misjudgments caused by spatial misalignment or content changes, and improving the reliability and automation level of the image verification process.
[0087] In an exemplary embodiment, identifying a start frame image and an end frame image corresponding to a scene under test in a corrected image sequence includes: performing the following first identification operation on each corrected image in the corrected image sequence as a first image to be identified, until a start frame image is identified: determining the feature similarity between a feature region of the first image to be identified and a feature region of a specified reference image to obtain a first feature similarity, wherein the specified reference image is the image to be verified or the first corrected image in the corrected image sequence; if the first feature similarity is less than a second similarity threshold, determining the first image to be identified as the start frame image; performing the following second identification operation on each corrected image after the start frame image in the corrected image sequence as a second image to be identified, until an end frame image is identified: determining the feature similarity between a feature region of the second image to be identified and a feature region of the start frame image to obtain a second feature similarity; if the second feature similarity is less than a third similarity threshold, determining the previous corrected image of the second image to be identified as the end frame image.
[0088] In related technologies, there is a problem of inaccurate recognition of start frame images / end frame images. To solve the above technical problem, this embodiment proposes a frame determination scheme based on feature similarity, which significantly improves the recognition accuracy of start frame images / end frame images, greatly reduces performance timing errors, and has a significant improvement in accuracy compared to related technologies. The test results can be directly used for performance acceptance.
[0089] In this embodiment, the feature region of the first image to be identified refers to a local visual region in the corrected image that is strongly correlated with the function of the test target. For example, the feature region of the first image to be identified may be an APP launch icon, a loading progress bar, or an interface background block. The feature region of the starting frame image refers to a local visual region in the starting frame image that is strongly correlated with the function of the test target. For example, the feature region of the starting frame image may be an APP launch icon, a loading progress bar, or an interface background block. The feature region of the second image to be identified refers to a local visual region in the corrected image following the starting frame image that is strongly correlated with the function of the test target. For example, the feature region of the second image to be identified may be an APP launch icon, a loading progress bar, or an interface background block. The feature region of the specified reference image refers to a local visual region in the specified reference image that is strongly correlated with the function of the test target. For example, the feature region of the specified reference image may be an APP launch icon, a loading progress bar, or an interface background block.
[0090] After obtaining the corrected image sequence, each corrected image can be selected frame by frame in chronological order as the first image to be identified, and a feature similarity comparison operation can be performed on the first image to be identified until the starting frame determination condition is met.
[0091] The specified reference image is either the image to be verified or the first corrected image in the sequence of corrected images. It can be understood that the specified reference image can be a standardized benchmark image retained after parameter validity verification (i.e., the image to be verified, which is the image obtained and verified during the test initialization phase); or it can be the earliest frame of corrected image acquired in the sequence of corrected images (i.e., the first frame image). It can be used as a dynamic reference when no benchmark image is available, so that the system can selectively rely on the pre-stored benchmark or the real-time first frame, thereby enhancing scene adaptability.
[0092] The second similarity threshold is a preset numerical threshold used to determine the starting frame, which can be used to distinguish the state boundary between "unchanged" images and "images that have begun to change".
[0093] Optionally, visual feature regions (such as APP icons, loading interface areas, etc.) related to the test scene are extracted from the first image to be identified, and the similarity between the feature regions of the first image to be identified and the corresponding feature regions in the specified reference image is calculated. For example, the similarity can be obtained based on the matching metric of image feature vectors (such as SIFT, ORB or HOG, etc.), and the numerical range can be [0,1], with the closer to 1 indicating greater similarity.
[0094] If the feature similarity between the feature region of the first image to be identified in the corrected image sequence and the feature region of the specified reference image is lower than the preset second similarity threshold (e.g., 0.75), it indicates that a significant visual change has occurred in the current screen, that is, the interface under test (such as the APP interface) has started to respond. At this time, the first image to be identified is marked as the starting frame.
[0095] In some embodiments, the interface element feature regions in the first image to be identified are extracted. For example, the interface element feature regions may include the SIFT keypoint and color histogram combination features of the status bar, navigation bar, and main icon area. The Euclidean distance and cosine similarity between the feature regions of the first image to be identified and the feature regions corresponding to the specified reference image are calculated respectively. The weighted fusion value of the Euclidean distance and cosine similarity between the feature regions of the first image to be identified and the feature regions corresponding to the specified reference image is used as the first feature similarity. When the first feature similarity is lower than the second similarity threshold (e.g., 0.72), and the image brightness and color distribution show obvious changes (e.g., icon highlighting, loading animation appears), the current first image to be identified is determined to be the starting frame image.
[0096] In some embodiments, semantic segmentation is performed on the interface layout area in the first image to be identified, and morphological feature vectors of button positions, text areas, and background color blocks are extracted; the morphological feature vectors are matched with the feature vectors of the corresponding feature areas in a specified reference image, and the matching results are normalized to obtain a first feature similarity; if the first feature similarity of the first image to be identified is lower than a second similarity threshold (e.g., 0.68) and a new interface element appears in the image (e.g., the "Application Management" label is highlighted), then the first image to be identified is identified as the starting frame image.
[0097] After the initial frame is identified, each subsequent corrected image is selected in chronological order as the second image to be identified, and subsequent feature comparison operations are performed until the end frame determination condition is met.
[0098] When the feature similarity between the second image to be identified and the starting frame is lower than the third similarity threshold (e.g., 0.65), it indicates that the current image is significantly different from the initial state and has stabilized, signifying that loading or switching has been completed. At this point, the previous frame of the second image to be identified (i.e., the last frame that still maintains a high similarity with the starting frame) is determined as the end frame, which can avoid recognition lag caused by inter-frame delay or caching. The third similarity threshold is a preset numerical threshold used to determine the end frame, which can be used to distinguish between the state boundaries of "still changing" and "having stabilized". Optionally, the third similarity threshold is less than the second similarity threshold.
[0099] It should be noted that the process of determining the feature similarity between the feature region of the second image to be identified and the feature region of the starting frame image can refer to the process of determining the feature similarity between the feature region of the first image to be identified and the feature region of the specified reference image, and will not be repeated here.
[0100] For example, by triggering automated commands (such as APP startup), the screen can be continuously captured and the region can be extracted and corrected frame by frame. The ROI of the image to be verified (Image A) after verification can be compared, and the starting frame (interface change) and ending frame (loading completed) can be determined according to the above-mentioned set criteria. Figure 3 This is a flowchart of an optional method for determining the start frame and the end frame according to an embodiment of this application, such as... Figure 3As shown, the process involves loading an image / ROI region, i.e., loading any image to be identified from the corrected image sequence (as the first or second image to be identified); cropping the ROI of the display screen area in the image, removing redundant background information, and obtaining a feature region containing only the display screen content; loading the lightweight backbone network MobileNetV3-Small, using MobileNetV3-Small as the base model for feature extraction, and loading pre-trained weights to ensure the effectiveness and generalization of feature extraction; and extracting shallow features by using the shallow convolutional layers of MobileNetV3-Small to process the ROI region of the corrected image. Feature extraction outputs a shallow feature map containing information such as texture, edge, and color distribution. Subsequently, feature vector pooling compression is performed, i.e., pooling operations (such as average pooling or max pooling) are used to reduce the dimensionality of the shallow feature map, compressing the multi-dimensional feature vector into a fixed-dimensional feature vector. Cosine similarity is calculated: for the first recognition operation, the cosine similarity between the compressed feature vector of the first image to be recognized and the feature vector of a specified reference image is calculated to obtain the first feature similarity; for the second recognition operation, the cosine similarity between the compressed feature vector of the second image to be recognized and the feature vector of the already recognized starting frame image is calculated to obtain the second feature similarity. Subsequently, the similarity value / batch comparison is output, that is, the first feature similarity and the second feature similarity obtained in a single calculation are output. If batch processing of the corrected image sequence is required, the similarity calculation of all images to be compared is completed in sequence, and the batch comparison results are output. The results are output, that is, if the first feature similarity is less than the second similarity threshold, the first image to be identified is determined as the starting frame image, and if the second feature similarity is less than the third similarity threshold, the previous corrected image of the second image to be identified is determined as the ending frame image.
[0101] Thus, the frame determination scheme based on feature similarity significantly improves the accuracy of start / end frame recognition, greatly reduces performance timing errors, and shows a significant improvement in accuracy compared to existing technologies. The test results can be directly used for performance acceptance, with high frame determination accuracy and reliable test results.
[0102] In this embodiment, each corrected image in the corrected image sequence is treated as the first image to be identified and a first identification operation is performed until the starting frame image is identified. Each corrected image after the starting frame image in the corrected image sequence is treated as the second image to be identified and a second identification operation is performed until the ending frame image is identified. This enables the dynamic determination of the starting and ending frames of the tested scene in the corrected image sequence based on feature similarity. It does not rely on fixed ROI parameters or manual intervention. The state transition moment can be accurately captured only through the relative change of visual features, thereby improving the robustness of frame recognition and the accuracy of timestamps.
[0103] In one exemplary embodiment, extracting the display area where the display screen is located in each test frame image in the test frame image sequence includes: identifying the geometric boundary features of the display screen in each test frame image to obtain the display area in each test frame image, wherein the geometric boundary features are used to identify the location area where the display screen is located.
[0104] In related technologies, some existing solutions, when extracting the display screen area from images captured by image acquisition devices (such as cameras), only use simple boundary recognition methods to directly perform region matching and parameter mapping. This can easily lead to deviations in the location of the region of interest, further affecting the accuracy of various test results. To solve the above problems, this method identifies the geometric boundary features of the display screen in each test frame image and extracts its corresponding region. This enables automatic and more accurate positioning of the display screen in the images captured by the image acquisition device, improving the accuracy of the test results.
[0105] In this embodiment, the geometric boundary features of the display screen in the test frame image refer to the shape, contour or edge structure features that can be extracted from the display screen in the test frame image. Optionally, the geometric boundary features may include, but are not limited to, straight line segments, corner points, closed contours, edge gradient direction and intensity distribution, etc. For example, the geometric boundary features may refer to the rectangular or quadrilateral outer contour of the display screen and the positions of its four corner points in the captured image, which can be used to characterize the spatial position and shape of the display screen in the image coordinate system.
[0106] Optionally, each test frame image in the test frame image sequence can be identified and extracted using image processing algorithms (such as HSV color space segmentation and contour detection) to identify and extract regions in each test frame image that conform to the optical characteristics of the display screen. By extracting pixel regions with high brightness, uniform color gamut, and rectangular or near-rectangular borders in each test frame image, and combining edge gradient changes, the physical geometric boundaries of the display screen in each test frame image are located and identified. After identifying the geometric boundaries of the display screen, the closed area enclosed by the boundaries is output as the effective pixel area of the display screen, forming a two-dimensional pixel mask, thereby obtaining the display screen area in each test frame image.
[0107] For example, each test frame image in the test frame image sequence can identify the geometric boundaries of the display screen through edge detection and rectangle fitting algorithms. First, Gaussian filtering is applied to each test frame image in the test frame image sequence to reduce noise. Then, the Canny edge detection operator is used to extract strong edge points in the image. Subsequently, Hough transform is used to detect candidate line segments, and the four most significant and approximately perpendicularly intersecting lines are selected as the four sides of the display screen. Finally, the coordinates of the four corner points of the display screen are calculated through the intersection points of the four sides, and the quadrilateral region is used as the geometric boundary of the display screen, thus obtaining the display screen area in each test frame image.
[0108] This embodiment identifies the geometric boundary features of the display screen in each test frame image and extracts its corresponding region, enabling automatic and more accurate positioning of the display screen in the image captured by the image acquisition device. This effectively suppresses background interference and non-target areas from interfering with subsequent processing, providing a stable and reliable input area for subsequent distortion correction and parameter mapping, and improving the robustness and consistency of display screen region extraction.
[0109] In an exemplary embodiment, identifying the geometric boundary features of the display screen in each test frame image includes: performing color space conversion on each test frame image to obtain each converted test frame image; and identifying N corner points of the display screen in each converted test frame image based on a preset display screen color range, wherein the geometric boundary features include N corner points, and N is a positive integer greater than or equal to 2.
[0110] In this embodiment, color space conversion refers to the mathematical transformation process of converting an image from one color representation system (such as RGB) to another color representation system (such as HSV). A color space can include the HSV color space, which is a three-dimensional color model composed of hue, saturation, and value. Hue represents the type of color (such as red, green, and blue), saturation represents the purity of the color, and value represents the lightness or brightness.
[0111] Optionally, a color space conversion operation is performed on each test frame image (usually in RGB color mode) acquired by the image acquisition device to convert each test frame image from RGB space to HSV (hue, saturation, brightness) color space. This can enhance the color feature expression capability of the display area in each test frame image, and make the inherent colors of the display (such as black borders, dark backlight areas, etc.) have more significant and stable distribution characteristics in the HSV space, thereby improving the accuracy of subsequent color-based region segmentation.
[0112] A corner point is a pixel in an image whose grayscale or color falls in two orthogonal directions. Optionally, a corner point can correspond to the corner position of an object's edge. For example, a corner point can refer to the geometric vertex of the physical bezel of a display screen, and can be used to characterize the planar boundary shape of the display screen.
[0113] Optionally, in the HSV color space of each test frame image, a pixel-level color mask is generated according to the preset display color range (e.g., H: 10°–30°, S: 40%–80%, V: 70%–95%). Then, the only region with the largest area and an aspect ratio close to the typical value of the cockpit display (e.g., 16:9) is selected through connected component analysis. Subsequently, the contour of this region is approximated by polygons (ε=0.02× contour perimeter), and the vertices of the approximated quadrilaterals are extracted as corner points. If there are fewer than four corner points (e.g., 2 corner points), the extreme points of the contour curvature are combined to complete the corners. Finally, the corner points are sorted in a clockwise direction to output N=4 stable corner points, ensuring that the geometric boundary of the display can still be accurately located under changes in illumination and local reflection interference.
[0114] For example, Figure 4 This is a flowchart of an optional corner recognition method according to an embodiment of this application, such as... Figure 4 As shown, in the preprocessing stage, each test frame image is converted to a different color space to obtain a converted test frame image. Based on the preset display screen color range, each converted test frame image is filtered by color threshold and grayscaled to obtain a binarized image. In the contour extraction stage, high-threshold binarization is used to strengthen the outer contour edge of the display screen. The outer contour is extracted from the binarized image, and the contour is smoothed by morphological closing operation to eliminate minor depressions. In the corner fitting stage, the contour of the display screen area is processed by convex hull processing. A polygon approximation algorithm is used to fit the smallest bounding polygon. The corner position is iteratively optimized to make the corner point completely aligned with the actual geometric boundary of the display screen. The result is output, that is, the coordinates of the N corner points of the display screen that are identified.
[0115] Through this embodiment, by converting color space and identifying corner points based on a preset color range, the geometric boundary features (corner points) of the display screen can be stably extracted from each frame of test image. This effectively suppresses the interference of ambient light changes, specular reflections, and non-uniform shadows on boundary detection, and improves the robustness and accuracy of display screen area positioning in complex environments.
[0116] In one exemplary embodiment, distortion correction is performed on the display area in each test frame image to obtain a corrected image sequence, including: performing distortion correction on the display area in each test frame image through perspective transformation to obtain a corrected image sequence.
[0117] In this embodiment, perspective transformation is an image transformation method based on projection geometry, which can be used to correct quadrilateral distortions (such as trapezoids and parallelograms) caused by the tilt of the shooting angle. By establishing a point-to-point nonlinear mapping relationship between the source image and the target image, geometric correction of the image is achieved.
[0118] Optionally, the coordinates of the four corner points of the display screen are identified from a test frame image as the source points of the perspective transformation. Then, based on a preset ideal rectangular target area (such as a rectangular box with the same resolution as the target reference image), the coordinates of the corresponding four target points are determined. Next, by calculating the perspective transformation matrix, the source point coordinates are mapped to the target point coordinates. All pixels in the display screen area of the test frame image are resampled and spatially transformed to generate a geometrically standardized image (i.e., a corrected image). The above process is repeated for each test frame image to form a sequence of corrected images.
[0119] For example, Figure 5 This is a flowchart of an optional perspective transformation according to an embodiment of this application, such as... Figure 5 As shown, the input parameters (i.e., the corner coordinates obtained above) are used; a dual-end coordinate system is defined, using the N corner coordinates of the display screen in the test frame image as the source points (source coordinates) of the perspective transformation, representing the outline of the display screen under the current distortion state. Based on the preset standard display screen size (such as physical resolution or target correction size), the corresponding regular rectangular target points (target coordinates) are defined (e.g., the four corner points correspond to the four vertices of the standard rectangle); based on the correspondence between the source coordinates and the target coordinates, the perspective transformation matrix (forward transformation matrix) is solved, and its inverse transformation matrix is calculated for subsequent coordinate inverse mapping. The original screen image is processed, the test frame image to be corrected is loaded, and the perspective transformation calculation is performed on the test frame image based on the forward transformation matrix. The display screen area is projected from the original distorted viewpoint to the target standard viewpoint, and the corrected display screen area is mapped back to the original test frame image's shooting size, maintaining consistent overall image resolution and proportion; the corrected display screen area is extracted and saved, and the corrected display screen area is extracted from the transformed image and saved as a corrected image sequence in chronological order, returning the result (i.e., outputting the corrected image sequence).
[0120] In this way, by using the corner point extraction and perspective transformation correction technology, we can effectively deal with screen tilt, slight camera offset and common image distortion within the normal range. After standardization, the image deviation is effectively controlled within an extremely low range, resulting in high frame determination accuracy and reliable test results. Furthermore, by adopting the correction approach of "corner point positioning + perspective transformation", we can specifically solve the image distortion problem caused by changes in environmental position such as camera offset and screen tilt, achieve accurate correction of distorted areas, and ensure the stability of screen area recognition.
[0121] In an optional embodiment, Figure 6 This is an optional test flowchart according to an embodiment of this application, such as... Figure 6As shown, the test initialization is performed, pre-storing the baseline image / ROI / automatic sequence; acquiring the original frame image and retrieving the B image; performing region extraction and distortion correction, segmenting the effective region and performing perspective transformation to generate the A image; confirming the validity of the ROI parameters and comparing feature fit; determining whether the test meets the standards, if not, prompting that the parameters have achieved the blocking test; if the test meets the standards, performing frame capture and judgment, triggering instructions, performing frame-by-frame correction and ROI comparison, and determining the start and end frame images; data calculation and storage, calculating performance indicators, and automatic storage; single-scene test completed.
[0122] Thus, through the above embodiments, it is not limited to APP startup time testing, but can adapt to various visual performance testing scenarios such as interface switching speed and pop-up response time. It is compatible with cockpit displays of different sizes and various cockpit APPs, without the need for customized development for specific scenarios. It can directly connect to existing cockpit testing platforms, has strong versatility, and is suitable for multiple performance testing scenarios. Furthermore, the core algorithms (perspective transformation, feature comparison model) are all mature industrial-grade technologies, requiring no special hardware customization and reusing existing test cameras and central control equipment. The system has excellent continuous operation stability, and the frame processing latency can meet the real-time requirements of high-performance testing, making it easy to promote and apply on a large scale. It has good engineering feasibility and excellent stability.
[0123] In this embodiment, perspective transformation is used to correct the distortion of the display screen area in each test frame image. This can transform non-rectangular distorted images caused by camera offset or display screen tilt into standard rectangular images, significantly improving the geometric consistency of the display screen area. This provides structurally regular input data for subsequent region matching and feature comparison, thereby enhancing the stability and repeatability of image processing.
[0124] Taking a smart cockpit as an example, the above embodiments can improve a visual testing system and method adapted to cockpit performance testing. This involves calibration, correction, and parameter verification technologies for the camera shooting area of cockpit products. It is applicable to cockpit function testing and interface response smoothness performance testing. The system can solve the aforementioned technical problems through a closed-loop process of "real-time extraction of test frames - distortion correction - dynamic verification of ROI parameters - frame state determination." The system includes an image acquisition module, a region extraction and correction module, a parameter verification module, and a parameter storage module. These modules work together to process each frame and determine its state during performance testing without manual intervention. It adapts to various visual performance testing scenarios, such as APP startup time and interface response speed, achieving universal adaptation to various cockpit visual testing scenarios and ensuring test stability and result reliability.
[0125] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0126] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as read-only memory (ROM) / random access memory (RAM), magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0127] According to another aspect of the embodiments of this application, a vision-based device testing apparatus is also provided. This vision-based device testing apparatus can be used to implement the vision-based device testing method provided in the above embodiments, and details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that implements a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0128] Figure 7 This is a structural block diagram of an optional vision-based device testing apparatus according to an embodiment of this application, such as... Figure 7 As shown, the vision-based device testing apparatus includes an acquisition unit 702, a correction unit 704, and a determination unit 706.
[0129] The acquisition unit 702 is used to acquire a sequence of test frame images corresponding to the interface under test displayed on the screen of the device under test during the testing process. The sequence of test frame images is obtained by continuously acquiring images of the interface under test through an image acquisition device.
[0130] The correction unit 704 is used to extract the display area where the display screen is located in each test frame image in the test frame image sequence, and to perform distortion correction on the display area in each test frame image to obtain a corrected image sequence, wherein the corrected image sequence includes the corrected image corresponding to each test frame image.
[0131] The determining unit 706 is used to identify the start frame image and the end frame image corresponding to the scene under test in the corrected image sequence, and to determine the test result of the scene under test based on the timestamp information of the start frame image and the timestamp information of the end frame image.
[0132] It should be noted that the acquisition unit 702 in this embodiment can be used to perform the above step S202, the correction unit 704 in this embodiment can be used to perform the above step S204, and the determination unit 706 in this embodiment can be used to perform the above step S206.
[0133] Through the embodiments provided in this application, during the testing of the device under test, a sequence of test frame images corresponding to the interface under test on the display screen of the device under test is acquired, and the display screen area in each test frame image is extracted and distortion corrected to obtain a corrected image sequence. This can eliminate the image distortion caused by the positional offset of the image acquisition device, the tilt of the display screen, or optical distortion, thereby ensuring that the geometric structure of the display screen area in each frame image remains consistent. This allows the recognition process of the start frame image and the end frame image to be based on stable, distortion-free visual features, avoiding regional positioning offset or feature misjudgment caused by image deformation. Finally, the test result is directly calculated based on the timestamp of the start frame image and the timestamp of the end frame image, so that the accuracy of the test result is not affected by external acquisition environment disturbances, significantly improving the test accuracy based on visual recognition. Therefore, it can solve the technical problem of low test accuracy in related technologies.
[0134] In an exemplary embodiment, the above-described apparatus further includes a first execution unit, which is configured to acquire an original frame image corresponding to an initial interface displayed on the display screen of the device under test, wherein the original frame image is obtained by image acquisition of the initial interface by an image acquisition device; extract the display screen area where the display screen is located in the original frame image, and perform distortion correction on the display screen area in the original frame image to obtain an image to be verified; map the feature region parameters of a target reference image onto the image to be verified to obtain a verification reference image, wherein the target reference image is a reference image with pre-calibrated feature regions corresponding to the scene under test, and the feature region parameters are used to indicate the calibrated feature regions in the target reference image; perform image verification on the verification reference image based on the target reference image to obtain the image verification result of the image to be verified, wherein acquiring the test frame image sequence is performed when the image verification result is that the verification is passed.
[0135] In an exemplary embodiment, the pixel coordinates of the image to be verified correspond one-to-one with the pixel coordinates of the target reference image, and the resolution of the image to be verified is the same as the resolution of the target reference image; the first execution unit is configured to determine that the image verification result is verified as passed when the feature similarity between the verification reference image and the target reference image is greater than or equal to a first similarity threshold; and to determine that the image verification result is verified as failed when the feature similarity between the verification reference image and the target reference image is less than the first similarity threshold.
[0136] In an exemplary embodiment, the determining unit 706 is configured to perform the following first recognition operation on each corrected image in the corrected image sequence as a first image to be recognized, until a starting frame image is recognized: determining the feature similarity between the feature region of the first image to be recognized and the feature region of a specified reference image to obtain a first feature similarity, wherein the specified reference image is the image to be verified or the first corrected image in the corrected image sequence; if the first feature similarity is less than a second similarity threshold, determining the first image to be recognized as the starting frame image; and performing the following second recognition operation on each corrected image after the starting frame image in the corrected image sequence as a second image to be recognized, until an ending frame image is recognized: determining the feature similarity between the feature region of the second image to be recognized and the feature region of the starting frame image to obtain a second feature similarity; if the second feature similarity is less than a third similarity threshold, determining the previous corrected image of the second image to be recognized as the ending frame image.
[0137] In an exemplary embodiment, the correction unit 704 is used to identify the geometric boundary features of the display screen in each test frame image to obtain the display screen area in each test frame image, wherein the geometric boundary features are used to identify the location area of the display screen.
[0138] In an exemplary embodiment, the correction unit 704 is configured to perform color space conversion on each test frame image to obtain each converted test frame image; based on a preset display screen color range, it identifies N corner points of the display screen in each converted test frame image, wherein the geometric boundary features include N corner points, and N is a positive integer greater than or equal to 2.
[0139] In an exemplary embodiment, the correction unit 704 is used to perform distortion correction on the display screen area in each test frame image through perspective transformation to obtain a corrected image sequence.
[0140] It should be noted that the above modules can be implemented by software or hardware. For the latter, they can be implemented in the following ways, but are not limited to: all the above modules are located in the same processor; or, the above modules are located in different processors in any combination.
[0141] According to another aspect of the embodiments of this application, a computer-readable storage medium is provided, the computer-readable storage medium including a stored program, wherein the program executes the steps in any of the above method embodiments when it is run.
[0142] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as USB flash drives, ROMs, RAMs, portable hard drives, magnetic disks, or optical disks.
[0143] According to another aspect of the embodiments of this application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor is configured to perform the steps of any of the method embodiments described above via the computer program. In an exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0144] Specific examples in this embodiment can be found in the examples described in the above embodiments and exemplary implementations, and will not be repeated here.
[0145] According to another aspect of the embodiments of this application, a computer program product is also provided, comprising a computer program / instructions containing program code for performing the methods shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication section 809, and / or installed from a removable medium 811. When the computer program is executed by a central processing unit 801, it performs various functions provided in the embodiments of this application. The sequence numbers of the embodiments of this application above are merely descriptive and do not represent the superiority or inferiority of the embodiments.
[0146] Figure 8 A schematic block diagram of a computer system architecture for implementing embodiments of the present application is shown. Figure 8As shown, the computer system 800 includes a Central Processing Unit (CPU) 801, which can perform various appropriate actions and processes based on programs stored in ROM 802 or programs loaded into RAM 803 from storage section 808. Random access memory 803 also stores various programs and data required for system operation. The CPU 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.
[0147] The following components are connected to I / O interface 805: an input section 806 including a keyboard, mouse, etc.; an output section 807 including a cathode ray tube (CRT), liquid crystal display (LCD), and speakers, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card, such as a local area network card or modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to I / O interface 805 as needed. Removable media 811, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 810 as needed so that computer programs read from them can be installed into storage section 808 as needed.
[0148] Specifically, according to embodiments of this application, the processes described in the various method flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 809, and / or installed from removable medium 811. When the computer program is executed by central processing unit 801, it performs various functions defined in the system of this application.
[0149] It should be noted that, Figure 8 The computer system 800 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0150] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. They can be implemented using computer-executable program code, and thus can be stored in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those described herein, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.
[0151] The above are merely preferred embodiments of this application and are not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the protection scope of this application.
Claims
1. A vision-based device testing method, characterized by, include: During the testing of the device under test, a sequence of test frame images corresponding to the interface under test displayed on the screen of the device under test is acquired. The sequence of test frame images is obtained by continuously acquiring images of the interface under test through an image acquisition device. The display screen area where the display screen is located is extracted from each test frame image in the test frame image sequence, and the display screen area in each test frame image is subjected to distortion correction to obtain a corrected image sequence, wherein the corrected image sequence includes the corrected image corresponding to each test frame image; Identify the start frame image and end frame image corresponding to the scene under test in the corrected image sequence, and determine the test result of the scene under test based on the timestamp information of the start frame image and the timestamp information of the end frame image.
2. The method of claim 1, wherein, The method further includes: Acquire the original frame image corresponding to the initial interface displayed on the screen of the device under test, wherein the original frame image is obtained by the image acquisition device acquiring the image of the initial interface; Extract the display screen area where the display screen is located in the original frame image, and perform distortion correction on the display screen area in the original frame image to obtain the image to be verified; The feature region parameters of the target reference image are mapped onto the image to be verified to obtain a verification reference image. The target reference image is a reference image with pre-calibrated feature regions corresponding to the scene under test. The feature region parameters are used to indicate the calibrated feature regions in the target reference image. Image verification is performed on the verification reference image based on the target reference image to obtain the image verification result of the image to be verified. The acquisition of the test frame image sequence is performed when the image verification result is that the verification is passed.
3. The method according to claim 2, characterized in that, The pixel coordinates of the image to be verified correspond one-to-one with the pixel coordinates of the target reference image, and the resolution of the image to be verified is the same as that of the target reference image. The step of performing image verification on the verification reference image based on the target reference image to obtain the image verification result of the image to be verified includes: If the feature similarity between the verification reference image and the target benchmark image is greater than or equal to a first similarity threshold, the image verification result is determined to be a successful verification. If the feature similarity between the verification reference image and the target baseline image is less than the first similarity threshold, the image verification result is determined to be verification failure.
4. The method according to claim 2, characterized in that, The step of identifying the start and end frame images in the corrected image sequence that correspond to the scene under test includes: Each corrected image in the corrected image sequence is treated as a first image to be identified, and the following first identification operation is performed until the starting frame image is identified: the feature similarity between the feature region of the first image to be identified and the feature region of a specified reference image is determined to obtain a first feature similarity, wherein the specified reference image is the image to be verified or the first corrected image in the corrected image sequence; if the first feature similarity is less than a second similarity threshold, the first image to be identified is determined as the starting frame image; Each corrected image after the starting frame image in the corrected image sequence is treated as a second image to be identified, and the following second identification operation is performed until the ending frame image is identified: the feature similarity between the feature region of the second image to be identified and the feature region of the starting frame image is determined to obtain a second feature similarity; if the second feature similarity is less than a third similarity threshold, the previous corrected image of the second image to be identified is determined as the ending frame image.
5. The method according to any one of claims 1 to 4, characterized in that, The step of extracting the display area where the display screen is located in each test frame image of the test frame image sequence includes: The geometric boundary features of the display screen in each test frame image are identified to obtain the display screen region in each test frame image, wherein the geometric boundary features are used to identify the location region of the display screen.
6. The method according to claim 5, characterized in that, The identification of the geometric boundary features of the display screen in each test frame image includes: Each test frame image is subjected to color space conversion to obtain the converted test frame image; Based on a preset display screen color range, N corner points of the display screen are identified in each converted test frame image, wherein the geometric boundary features include the N corner points, and N is a positive integer greater than or equal to 2.
7. The method according to any one of claims 1 to 4, characterized in that, The step of performing distortion correction on the display screen area in each test frame image to obtain a corrected image sequence includes: The distortion of the display screen area in each test frame image is corrected by perspective transformation to obtain the corrected image sequence.
8. A vision-based device testing apparatus, characterized in that, include: The acquisition unit is used to acquire a sequence of test frame images corresponding to the interface under test displayed on the screen of the device under test during the testing process. The sequence of test frame images is obtained by continuously acquiring images of the interface under test through an image acquisition device. The correction unit is used to extract the display screen area where the display screen is located in each test frame image in the test frame image sequence, and to perform distortion correction on the display screen area in each test frame image to obtain a corrected image sequence, wherein the corrected image sequence includes a corrected image corresponding to each test frame image; The determining unit is used to identify the start frame image and the end frame image corresponding to the scene under test in the corrected image sequence, and to determine the test result of the scene under test based on the timestamp information of the start frame image and the timestamp information of the end frame image.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the method according to any one of claims 1 to 7.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.