Small model-based in-vehicle infotainment test method and system

Through lightweight recognition models and multimodal data analysis, the problems of manual dependence and insufficient automation tools in in-vehicle infotainment system testing are solved, and efficient and reliable vehicle test automation and report generation are achieved.

CN120743785AInactive Publication Date: 2025-10-03TIANJIN XIAOBO ZHILIAN INFORMATION TECHNOLOGY CO LTD

Patent Information

Application Number
CN202511213173.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-10-03
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing in-vehicle infotainment system testing relies on manual operations, which is inefficient and inconsistent. Traditional automated tools lack intelligent recognition capabilities and are difficult to adapt to different vehicle models and interface design changes. They also lack multimodal data analysis capabilities and cannot meet the quality assurance requirements of rapid iteration and large-scale deployment.

Method used

A lightweight recognition model is used to extract vehicle screen elements, and interface elements are identified through lightweight convolutional neural networks and sequence modeling algorithms. Automated test scripts are generated, and multi-dimensional data is collected by combining industrial cameras, microphones, and OBD interfaces for multimodal data analysis and test report generation.

Benefits of technology

It improves the automation level and result reliability of vehicle computer testing, realizes intelligent recognition and operation of different vehicle computer interfaces, generates standardized test reports, reduces manual maintenance costs, and improves test efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120743785A_ABST
    Figure CN120743785A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and discloses an in-vehicle testing method and system based on a small model. The method comprises the following steps: acquiring an in-vehicle machine screen image, and extracting interface elements through a lightweight recognition model to generate a recognition result; analyzing element function association to determine operation types to form a test operation sequence; an ADB instruction and a mechanical arm action are generated according to the operation sequence to construct a test script; executing the script to drive the vehicle machine to test and collecting multi-dimensional response data; and analyzing test data, calculating a passing rate and response time, and generating a verification report. According to the method, the technical problems of low interface element identification accuracy, poor test script adaptability and insufficient multi-modal data analysis capability in the traditional vehicle-mounted terminal test are solved, and the automation degree of the vehicle-mounted terminal test and the reliability of the test result are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a vehicle computer testing method and system based on a small model. Background Art

[0002] Existing in-vehicle infotainment system testing mainly relies on manual operation and traditional automation tools. Testers verify system functions by manually clicking on the vehicle screen, observing interface responses and recording test results. Some test scenarios use coordinate-based automated scripts or simple image recognition technology to simulate user operations. Test data collection is usually limited to single-dimensional information collection, such as only recording the success or failure of an operation or simple response time data. Test result analysis and report generation are mainly completed through manual collation and statistics.

[0003] However, existing technologies have obvious shortcomings. Manual testing is inefficient and easily affected by subjective factors, and cannot ensure the consistency and accuracy of test operations. Traditional automation tools lack the ability to intelligently identify vehicle interface elements and are difficult to adapt to changes in different vehicle models and interface designs. Test scripts based on coordinate positioning are prone to failure after the interface layout is adjusted and require frequent maintenance and updates. Single-dimensional data collection cannot fully reflect the actual operating status of the vehicle system. Test analysis lacks multimodal data fusion and deep mining capabilities. The report generation process is cumbersome and the degree of standardization is not high.

[0004] More importantly, with the increasing functional complexity of vehicle-mounted systems and the diversification of test scenarios, existing technologies face even more severe challenges. Traditional testing methods cannot effectively handle the problems of intelligent recognition of vehicle-mounted interface elements and understanding of operational intentions. They lack the ability to deploy lightweight models on edge devices, making it difficult to achieve real-time collaborative analysis of multimodal data. They cannot automatically generate standardized test execution scripts, and lack intelligent detection and adaptive adjustment mechanisms for abnormal conditions during the test process. These problems lead to inefficient and costly vehicle-mounted testing, making it difficult to meet the quality assurance requirements of rapid iteration and large-scale deployment of modern vehicle-mounted systems. Summary of the Invention

[0005] This application provides a vehicle computer testing method and system based on a small model, which is used to solve the technical problems of low interface element recognition accuracy, poor test script adaptability, and insufficient multimodal data analysis capabilities in traditional vehicle computer testing, significantly improving the degree of automation of vehicle computer testing and the reliability of test results.

[0006] In a first aspect, the present application provides a vehicle computer testing method based on a small model, the vehicle computer testing method based on a small model comprising:

[0007] Step S1: Obtain the vehicle screen display image, extract the buttons, icons and text elements on the screen through a lightweight recognition model, and generate an interface element recognition result including coordinate position and element type;

[0008] Step S2: Analyzing the functional association and operation sequence between elements based on the interface element recognition results, determining the touch click, slide switch and menu navigation operation types, and obtaining the vehicle computer test operation sequence;

[0009] Step S3: Generate ADB touch screen control instructions and robotic arm execution actions according to the operation type and coordinate parameters in the vehicle computer test operation sequence, and build an automated test execution script;

[0010] Step S4: running the automated test execution script to drive the vehicle system to execute the test case, while collecting vehicle response data through the industrial camera, microphone and OBD interface to obtain a multi-dimensional test data set;

[0011] Step S5: Analyze the interface changes, audio output and bus communication status in the multi-dimensional test data set, calculate the test pass rate and response time indicators, and generate a vehicle computer function verification report.

[0012] In a second aspect, the present application provides a vehicle computer test system based on a small model, the vehicle computer test system based on a small model comprising:

[0013] The extraction module is used to obtain the vehicle screen display image, extract the buttons, icons and text elements on the screen through a lightweight recognition model, and generate interface element recognition results including coordinate positions and element types;

[0014] An analysis module is used to analyze the functional association and operation sequence between elements based on the interface element recognition results, determine the touch click, slide switch and menu navigation operation types, and obtain the vehicle computer test operation sequence;

[0015] A generation module is used to generate ADB touch screen control instructions and robotic arm execution actions according to the operation type and coordinate parameters in the vehicle computer test operation sequence, and build an automated test execution script;

[0016] An operation module is used to run the automated test execution script to drive the vehicle system to execute test cases, and at the same time collect vehicle response data through industrial cameras, microphones and OBD interfaces to obtain a multi-dimensional test data set;

[0017] The calculation module is used to analyze the interface changes, audio output and bus communication status in the multi-dimensional test data set, calculate the test pass rate and response time indicators, and generate a vehicle function verification report.

[0018] In a third aspect, a vehicle-computer test device based on a small model is provided, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory so that the vehicle-computer test device based on the small model executes the above-mentioned vehicle-computer test method based on the small model.

[0019] In a fourth aspect, a computer-readable storage medium is provided, wherein instructions are stored in the computer-readable storage medium, which, when executed on a computer, enables the computer to execute the above-mentioned small-model-based vehicle-machine testing method.

[0020] In the technical solution provided in this application, the vehicle screen image is intelligently analyzed through a lightweight recognition model, and the coordinate position and type information of buttons, icons and text elements are automatically extracted, which solves the problems of low accuracy and poor adaptability caused by the traditional testing method relying on manual recognition and fixed coordinate positioning. The interface element recognition results lay an accurate data foundation for the subsequent generation of test operation sequences. Based on the analysis and processing of functional associations and operation sequences between interface elements, the operation types such as touch clicks, sliding switches and menu navigation are intelligently determined, and the dependence of traditional test scripts on changes in interface layout is eliminated. The automatic generation of vehicle test operation sequences significantly reduces the labor cost of test script writing and maintenance. The coordinated cooperation of ADB touch screen control instructions and robotic arm execution actions realizes non-invasive automated test execution, avoiding the inconsistency and subjectivity of manual operations in traditional testing methods. The construction of automated test execution scripts makes the test process standardized and repeatable.

[0021] Multi-dimensional data collection from industrial cameras, microphones, and OBD interfaces overcomes the limitations of traditional single-dimensional test data. This multi-dimensional test data set comprehensively reflects the true operating status of the vehicle-mounted system by integrating visual, audio, and bus communication information. Comprehensive analysis and processing of interface changes, audio output, and bus communication status overcomes the inadequate detection capabilities of traditional testing methods for system anomalies. Automatic calculation of test pass rates and response time metrics eliminates the errors and inefficiencies of manual statistics. Automatic generation of vehicle-mounted system function verification reports revolutionizes the cumbersome and low-standardization processes of traditional report writing. In particular, the application of lightweight convolutional neural networks and sequence modeling algorithms in vehicle-mounted system testing enables the efficient execution of complex interface element recognition and operation intent understanding tasks on resource-constrained edge devices, avoiding the technical bottlenecks of traditional deep learning models with large parameters and slow inference speed. The introduction of knowledge distillation technology and attention mechanisms further ensures that small models maintain high-precision recognition performance while maintaining their lightweight characteristics. The application of multi-head attention mechanisms in operation sequence generation enables test scripts to intelligently understand the complex relationships between interface elements, significantly improving their versatility and robustness across diverse vehicle-mounted system interface designs. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0023] Figure 1 A schematic diagram of an embodiment of a vehicle computer testing method based on a small model in an embodiment of the present application;

[0024] Figure 2 A schematic diagram of an embodiment of a vehicle-mounted computer testing system based on a small model in an embodiment of the present application;

[0025] Figure 3 It is a schematic block diagram of the structure of a vehicle computer testing device based on a small model in an embodiment of the present invention. DETAILED DESCRIPTION

[0026] An embodiment of the present application provides a vehicle-computer testing method and system based on a small model. The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0027] For ease of understanding, the specific process of the embodiment of the present application is described below. Figure 1 In the embodiment of the present application, an embodiment of the vehicle computer testing method based on a small model includes:

[0028] Step S1: Obtain the vehicle screen display image, extract the buttons, icons and text elements on the screen through a lightweight recognition model, and generate an interface element recognition result including coordinate position and element type;

[0029] Step S2: Analyze the functional association and operation sequence between elements based on the interface element recognition results, determine the touch click, slide switch and menu navigation operation types, and obtain the vehicle computer test operation sequence;

[0030] Step S3: Generate ADB touch screen control instructions and robotic arm execution actions based on the operation type and coordinate parameters in the vehicle computer test operation sequence, and build an automated test execution script;

[0031] Step S4: Run the automated test execution script to drive the vehicle system to execute the test case, and at the same time collect the vehicle system response data through the industrial camera, microphone and OBD interface to obtain a multi-dimensional test data set;

[0032] Step S5: Analyze the interface changes, audio output, and bus communication status in the multi-dimensional test data set, calculate the test pass rate and response time indicators, and generate a vehicle computer function verification report.

[0033] It is understandable that the execution subject of this application can be a vehicle-mounted test system based on a small model, or a terminal or a server, which is not limited here. The embodiment of this application is described by taking the server as the execution subject as an example.

[0034] Specifically, the image data is converted from the original output format of the vehicle computer to a standard RGB format, and images of different resolutions are uniformly normalized to a fixed size to obtain a standardized vehicle computer screen image. The lightweight recognition model uses the MobileNetV3 architecture. This architecture decomposes the standard convolution into two steps: depthwise convolution and pointwise convolution through depthwise separable convolution. Depthwise convolution performs independent spatial filtering on each input channel, while pointwise convolution combines information from different channels. The linear bottleneck structure performs nonlinear transformations in low-dimensional space and then maps it back to high-dimensional space. The channel attention mechanism compresses the spatial feature map into a single value through global average pooling, and then generates channel weights through two fully connected layers. Knowledge distillation technology uses the softmax function to calculate the probability distribution of the teacher network output as soft labels, guiding the student network to learn richer feature representations. The post-processing stage performs non-maximum suppression on the detection results, and filters low-quality detection boxes based on the confidence threshold. The final output is the interface element recognition result, which includes the bounding box coordinates, confidence score, and category label.

[0035] Based on the results of interface element recognition, the lightweight sequence modeling network converts the coordinate position and element type into a fixed-dimensional vector representation. The position encoding uses sine and cosine functions to map the two-dimensional coordinates into high-dimensional vectors. The temporal modeling processes the element sequence through the Transformer encoder. The multi-head attention mechanism calculates the similarity between the query vector, key vector and value vector, and calculates the attention weight through the scaled dot product attention formula. The element association matrix is ​​generated to reflect the spatial and functional relationship between the interface elements. The grouping clustering algorithm classifies similar elements into the same functional module according to the association threshold, identifies the menu navigation path and the interaction logic chain, and the operation sequence generator converts the abstract functional relationship into specific operation instructions according to the node connection relationship in the functional operation path diagram, including the target coordinates of the click operation, the start and end positions of the sliding operation, and the duration of the long press operation.

[0036] Each operation type in the vehicle-computer test operation sequence corresponds to a specific ADB command template. The click operation matches the "input tap" template, and the sliding operation matches the "input swipe" template. The system automatically selects the corresponding template and fills in the coordinate parameters according to the operation type. The robot arm execution action is calculated by the inverse kinematics algorithm. According to the target screen coordinates and the current position of the robot arm, the angle of rotation of each joint is calculated. The spatial coordinate conversion maps the two-dimensional screen coordinates to the three-dimensional robot arm workspace coordinates. The script arrangement module arranges instructions according to the execution order of the test case, inserts appropriate waiting time between consecutive operations, and synchronizes multiple device actions that need to be executed simultaneously. Syntax checking verifies the correctness of the generated script syntax. Logical verification ensures that the operation sequence conforms to the interactive logic of the vehicle-computer interface. The exception handling mechanism automatically retries when an operation failure is detected. The retry mechanism sets the maximum number of retries and the retry interval.

[0037] After receiving the test execution request, the Jenkins task scheduling system sends the automated test execution script to the edge device through the network interface. The edge device parses the script content and starts the corresponding test module. The industrial camera uses high frame rate mode to record the car screen in real time. The microphone collects the car audio output at a fixed sampling rate. The OBD interface monitors the CAN bus communication frames. The timestamp alignment algorithm synchronizes data from different sources according to a unified time base. The visual data contains the timestamp and pixel information of each frame of the image, the audio data contains the timestamp and amplitude value of the sampling point, and the bus data contains the timestamp and data content of the CAN frame. Data format standardization converts raw data of different formats into a unified data structure. The feature extraction algorithm extracts key features from multimodal data. The anomaly detection algorithm identifies abnormal events during the test process through threshold comparison and pattern matching.

[0038] Frame-by-frame image difference calculation obtains the changed area by subtracting the pixel values ​​of adjacent frames. The pixel change rate is calculated by dividing the number of changed pixels by the total number of pixels. The interface response time is obtained by detecting the difference between the time when the operation instruction is sent and the start time of the interface change. The short-time Fourier transform converts the audio signal from the time domain to the frequency domain, extracts the spectral features and calculates the similarity with the expected audio template. CAN frame parsing decodes the meaning of each signal according to a predefined data dictionary. The body control signals include vehicle speed, steering angle and braking status. The fault code is read from the vehicle control unit through the diagnostic communication protocol. The weighted algorithm assigns weights according to the importance of different test indicators, and calculates the comprehensive test indicators through weighted summation.

[0039] The test pass rate is calculated by dividing the number of successfully executed test cases by the total number of test cases. The response time indicator is obtained by averaging the response time of all test operations. The BERT text generation model converts structured test data into natural language descriptions. The preset template provides the basic framework and format specifications of the report. The test result description text includes a test summary, detailed results, and problem analysis. The formatting layout arranges the text content in a standard format. The pass rate pie chart shows the proportional relationship between successful and failed cases. The response time trend chart shows the changes in response time during the test. The defect distribution statistics table classifies and counts the defects according to defect type and severity. For example, during a test of an in-vehicle music application, the lightweight recognition model detected the coordinates of the play button on the music interface as (512, 384). Based on functional association analysis, the operation sequence generator determined that it was necessary to first click the music icon to enter the application and then click the play button. The generated ADB commands were "input tap 256 200" and "input tap 512 384." The robotic arm converted the screen coordinates into joint angles based on inverse kinematics calculations. During the test, an industrial camera captured interface changes, and frame difference analysis detected that the interface began to change 200 milliseconds after the play button was clicked. The microphone captured the audio output, and spectrum analysis confirmed that the music was playing properly. Finally, a comprehensive test report was generated that included interface response time, audio quality assessment, and functional correctness verification.

[0040] In a specific embodiment, step S1 includes:

[0041] Capture the vehicle screen display image in real time through the HDMI interface of the edge device, convert the image data into standard RGB format and perform size normalization to obtain a standardized vehicle screen image;

[0042] A lightweight convolutional neural network based on the MobileNetV3 architecture is used to extract features from standardized vehicle screen images, generating image feature maps through depthwise separable convolution and linear bottleneck structures.

[0043] The channel attention mechanism is used to perform weighted processing on the image feature map, and the channel weight is calculated through global average pooling and fully connected layers to output the interface element features;

[0044] The soft label knowledge of the teacher network is transferred to the student network through knowledge distillation technology, and the probability distribution is calculated using the softmax function to generate the element classification probability;

[0045] The identified interface elements are post-processed to extract the bounding box coordinates, confidence score and category label of each element, and output structured interface element recognition result data.

[0046] Specifically, when the edge device captures the image displayed on the vehicle screen in real time through the HDMI interface, it obtains the original image data stream from the video output end of the vehicle system. This data stream usually uses the YUV color space format. The image processing chip built into the edge device transforms the color space of each pixel in the YUV format according to the conversion matrix. The Y component represents brightness information, and the U and V components represent chromaticity information. During the conversion process, the Y component is directly mapped to the brightness basis in RGB. The U component is calculated through linear transformation to obtain the difference between the blue component and the brightness, and the V component is calculated to obtain the difference between the red component and the brightness. Finally, the three YUV components are converted into three RGB components through matrix operation. The size normalization process is used to scale the input image to a fixed size to account for the screen resolution differences of different vehicle models. The scaling algorithm uses bilinear interpolation. For magnification, new pixels are inserted between the original pixels. The RGB value of the interpolated pixel is calculated by weighted average of the four surrounding pixels. For reduction, representative pixels are selected through sampling. The normalized standardized vehicle screen image has a uniform pixel size and color format.

[0047] After receiving the standardized car screen image, the lightweight convolutional neural network with the MobileNetV3 architecture performs a depth-wise separable convolution operation. Traditional convolution requires applying a set of convolution kernels to each channel of the input, while depth-wise separable convolution decomposes this process into two independent steps. In the depth-wise convolution stage, a convolution kernel is applied to each channel of the input image. The convolution kernel of each channel slides in the two-dimensional space of the channel, and the product and accumulation of the convolution kernel and the pixel value of the corresponding area are calculated to generate the feature map of the channel. In the point-by-point convolution stage, a 1x1 convolution kernel is used to perform a linear combination of the output of the depth-wise convolution between channels, and the feature information of different channels is fused to generate the final feature map. The linear bottleneck structure expands the number of channels of the input feature map to a higher dimension through 1x1 convolution, and then applies depth-wise separable convolution for feature extraction in the high-dimensional space. Finally, the number of channels is compressed back to the original dimension through 1x1 convolution. This structure significantly reduces the computational parameters while maintaining the feature expression capability.

[0048] When the channel attention mechanism performs weighted processing on the image feature map, the global average pooling operation averages each channel of the feature map in the spatial dimension. Assuming that the feature map size is H×W×C, where H is the height, W is the width, and C is the number of channels, global average pooling adds the H×W pixel values ​​of each channel and divides it by the total number of pixels HW to obtain the global feature value of the channel. After global average pooling, the feature map is compressed from the three-dimensional tensor H×W×C to a one-dimensional vector C, which is input into two fully connected layers. The first fully connected layer compresses the number of channels C to C / r, where r is the compression ratio. The second fully connected layer restores the dimension to C. The ReLU activation function is used between the two fully connected layers to introduce nonlinear transformation. Finally, the output value is compressed to between 0 and 1 through the Sigmoid function as the channel weight. The weight value of each channel is multiplied by all pixel values ​​of the corresponding channel of the original feature map to generate the weighted interface element features.

[0049] In the knowledge distillation technology, the logits vector output by the teacher network is converted into a probability distribution through the softmax function. The softmax function performs an exponential operation on the logits value of each category and then normalizes it by dividing it by the sum of the exponential values ​​of all categories. The temperature parameter T controls the smoothness of the probability distribution. When T is greater than 1, the probability distribution becomes smoother and contains more similarity information between categories. The student network learns knowledge by minimizing the cross-entropy loss with the soft label of the teacher network. The cross-entropy loss calculates the difference between the probability distribution of the teacher network and the probability distribution of the student network. The parameters of the student network are updated according to the loss gradient through the backpropagation algorithm, so that the output of the student network gradually approaches the soft label distribution of the teacher network. This knowledge transfer process enables the lightweight student network to obtain recognition capabilities close to those of the large teacher network.

[0050] In the interface element post-processing stage, the detection results output by the network are confidence filtered, and detection frames with confidence higher than the preset threshold are retained. Then, the non-maximum suppression algorithm is applied to eliminate duplicate detections. Non-maximum suppression calculates the ratio of the overlapping area of ​​each two detection frames to the union area as the overlap. When the overlap exceeds the threshold, the detection frame with higher confidence is retained and the detection frame with lower confidence is deleted. The bounding box coordinate extraction includes the pixel coordinates of the upper left and lower right corners of the detection frame. The category label corresponds to the specific element type in the vehicle interface, such as button, icon or text. The confidence score reflects the network's confidence in the detection result. The structured interface element recognition result data organizes all detected element information into a unified data format. For example, in the car navigation interface test, the edge device captures the navigation screen with a resolution of 1920×1080 through the HDMI interface. The image conversion module scales the RGB image to 608×608. The MobileNetV3 network extracts multi-layer features and then uses the channel attention mechanism to enhance the feature representation of the map area and control buttons. During the knowledge distillation training process, the teacher network identifies elements such as the map display area, zoom button, search box, and route planning button. The student network also recognizes these elements after learning the soft label distribution of the teacher network. In the post-processing stage, the bounding box coordinates of the search box are extracted as (100, 50, 300, 90), with a confidence of 0.92 and a category label of text input box. The coordinates of the zoom button are (500, 400, 550, 450), with a confidence of 0.87 and a category label of control button.

[0051] In a specific embodiment, step S2 includes:

[0052] The coordinate position and element type in the interface element recognition result are input into the lightweight sequence modeling network based on the Transformer architecture for encoding processing to obtain the semantic feature vector of the interface element;

[0053] Perform position encoding and temporal modeling on the semantic feature vectors of interface elements, calculate the spatial correlation weights between elements through a multi-head attention mechanism, and obtain the element correlation matrix;

[0054] The functional modules of the vehicle interface are grouped and clustered based on the element correlation matrix to identify the menu navigation path and interaction logic chain, and obtain the functional operation path diagram;

[0055] The operation sequence is generated and processed according to the node connection relationship in the functional operation path diagram, and the click, slide, and long press operation types are matched and combined with the target coordinates to obtain the vehicle computer test operation sequence.

[0056] Specifically, when the lightweight sequence modeling network receives the interface element recognition results based on the Transformer architecture, it converts the coordinate position and element type of each interface element into a numerical vector representation. The coordinate position includes the horizontal and vertical coordinate values ​​of the element on the car screen. The element type uses a one-hot encoding method to convert categories such as buttons, icons, and text into binary vectors. The encoding process concatenates the coordinate value and type vector to obtain an input vector of fixed length. This vector is mapped to the hidden dimension space of the model through a linear transformation layer to obtain the initial interface element semantic feature vector. The semantic feature vector contains the position information and functional attribute information of each interface element, laying the data foundation for subsequent sequence modeling.

[0057] The position encoding process integrates the two-dimensional spatial position information of the interface elements on the screen into the semantic feature vector. The horizontal and vertical coordinates are encoded by the sine and cosine functions respectively. The frequency parameters of the sine and cosine functions are dynamically adjusted according to the size of the coordinate values, so that elements in adjacent positions have similar position encodings, while elements that are farther away have significantly different encodings. The temporal modeling process considers the operation order relationship between the interface elements and arranges the elements into a sequence according to the visual scanning order from left to right and from top to bottom. The multi-head attention mechanism projects the input semantic feature vector into a query vector, a key vector and a value vector respectively. Each attention head independently calculates the similarity between the query vector and the key vector. The similarity is obtained through the vector dot product operation, which represents the association strength between the two interface elements. The similarity between all element pairs constitutes an element association matrix. Each value in the matrix reflects the spatial association weight of the corresponding element pair.

[0058] The grouping clustering process identifies interface elements with similar functions based on the element correlation matrix. The clustering algorithm classifies elements with a correlation higher than a threshold into the same functional module. For example, buttons related to music playback will be clustered into the music functional module, and navigation-related controls will be clustered into the navigation functional module. Menu navigation path recognition analyzes the jump relationship between different functional modules and identifies the path sequence for users to enter sub-functions from the main interface. The interactive logic chain construction takes into account the dependencies between interface elements. Some operations must be performed after other operations are completed. The functional operation path diagram organizes all identified functional modules and navigation paths in the form of a graph structure. The nodes in the graph represent interface elements or functional states, and the edges represent operation conversion relationships.

[0059] The operation sequence generation process determines the execution order of the test operations based on the node connection relationship in the functional operation path diagram. The click operation match is applicable to elements of button and icon types. The operation target coordinates directly use the center point coordinates of the element. The sliding operation match is applicable to list and page switching scenarios. The starting coordinates select the edge position of the sliding area, and the ending coordinates are calculated based on the sliding direction and distance. The long press operation match is applicable to elements that need to pop up a context menu. The target coordinates use the center point coordinates of the element. The operation duration is set according to the interface response characteristics. The matching combination process binds the operation type with the target coordinates to obtain the operation instruction. The vehicle test operation sequence arranges all operation instructions according to the execution order in the path diagram.

[0060] In a specific embodiment, step S3 includes:

[0061] Parse the operation type and coordinate parameters in the vehicle computer test operation sequence, match the preset ADB command template according to the operation type, perform command splicing processing, and obtain the ADB touch screen control command set;

[0062] The joint angles and motion trajectories of the robotic arm are calculated based on the coordinate parameters in the vehicle-mounted computer test operation sequence. The spatial coordinates are transformed using the inverse kinematics algorithm to obtain the robotic arm's execution motion instructions.

[0063] The ADB touch screen control instruction set and the robotic arm execution action instruction are scripted according to the time sequence relationship, and waiting time and synchronization identifiers are inserted to obtain a time-sequential test instruction sequence;

[0064] Perform syntax checking and logic verification on the sequence of sequential test instructions, add exception handling and retry mechanisms, and obtain an automated test execution script.

[0065] Specifically, the vehicle computer test operation sequence parsing process extracts the operation type identifier and target coordinate value in each operation instruction. The operation types include three basic interaction methods: click, slide, and long press. The coordinate parameters contain the pixel position information of the horizontal and vertical coordinates. The ADB command template matching selects the corresponding instruction format according to the operation type. The click operation matches the "adb shell input tap" template, the slide operation matches the "adb shell input swipe" template, and the long press operation matches the "adb shell input tap" template and adds a duration parameter. The instruction splicing process connects the command string corresponding to the operation type with the coordinate parameter. The coordinate parameters of the click operation are directly inserted into the horizontal and vertical coordinate positions of the tap command. The slide operation requires four parameters, namely the starting coordinate and the ending coordinate, to be inserted into the swipe command. The long press operation adds the duration in milliseconds after the tap command. After splicing is completed, the ADB touch screen control instruction is obtained. All generated instructions are organized into an ADB touch screen control instruction set in the order of the operation sequence.

[0066] The calculation of the robot arm joint angle is based on the spatial coordinate conversion of the coordinate parameters in the vehicle computer test operation sequence. The inverse kinematics algorithm converts the two-dimensional screen coordinates into the angle value of the robot arm joint space. The algorithm establishes a spatial mapping relationship between the robot arm end effector and the vehicle computer screen. The horizontal coordinate of the screen plane corresponds to the rotation angle of the robot arm base, and the vertical coordinate corresponds to the pitch angle combination of the robot arm joint. The inverse kinematics calculation process reversely derives the angle setting of each joint based on the target screen position. The base joint angle is calculated by the inverse tangent function of the screen horizontal coordinate and the robot arm working radius. The shoulder joint and elbow joint angles are obtained by solving the geometric constraint equations. The motion trajectory planning generates a smooth joint angle change curve between the starting position and the target position to avoid sudden changes and vibrations during the movement of the robot arm. The time interval of the trajectory points is determined according to the movement speed limit of the robot arm. Each trajectory point contains the target angle and arrival time of each joint. The robot arm execution action instruction contains the angle sequence and time series of all trajectory points.

[0067] The script arrangement processing combines the ADB touch screen control instruction set and the robotic arm execution action instruction according to the execution order of the test case. The timing relationship analysis determines which instructions need to be executed synchronously and which instructions need to be executed sequentially. The touch screen instruction can only be sent after the robotic arm moves to the target position. After the touch screen operation is completed, it is necessary to wait for the vehicle interface to respond. The waiting time is inserted to set a fixed delay according to the response characteristics of the vehicle interface. The waiting time after general operations is set to a shorter time interval. The waiting time after complex operations such as application startup is set to a longer time interval. The synchronization identifier marks the multi-device actions that need to be coordinated. When the robotic arm reaches the specified position, the ADB instruction is triggered to be sent. The timed test instruction sequence arranges all instructions in order of execution time. Each instruction contains an execution timestamp, instruction type and parameter content.

[0068] The syntax check process verifies the grammatical correctness of each instruction in the timing test instruction sequence. The ADB instruction check includes the verification of the command keyword spelling, the number of parameters and the parameter format. The robotic arm instruction check includes the verification of the joint angle range and the movement speed limit. The logic verification process analyzes the execution logic rationality of the instruction sequence and checks whether there are invalid operation combinations or conflicting instruction sequences. The exception handling mechanism adds processing logic for possible error conditions that may occur during the execution process. When the ADB instruction fails to execute, the error message is automatically recorded and tried to be resent. When the robotic arm moves abnormally, the current action is automatically stopped and returned to a safe position. The retry mechanism sets the maximum number of retries and the retry interval. The retry number limit for a single instruction prevents infinite loops. The retry interval gives the vehicle interface sufficient recovery time. The automated test execution script integrates all instructions, waiting time, synchronization control and exception handling logic.

[0069] In a specific embodiment, step S4 includes:

[0070] The automated test execution script is triggered and executed through the Jenkins task scheduling system, test instructions are sent to the edge device, and the vehicle system test is started to obtain test execution status feedback;

[0071] Based on the test execution status feedback, the industrial camera, microphone and OBD interface are synchronously started to perform multimodal data acquisition and processing, and the vehicle screen image, audio output and CAN bus communication frames are recorded in real time to obtain the original test data stream;

[0072] Perform timestamp alignment and data format standardization on the original test data stream, synchronize and integrate visual data, audio data, and bus data according to a unified time base to obtain time-aligned multimodal data;

[0073] The time-aligned multimodal data is comprehensively analyzed and processed through feature extraction and anomaly detection algorithms to identify response delays and functional abnormalities during the test process and obtain a multidimensional test data set.

[0074] Specifically, after receiving the test task request, the Jenkins task scheduling system parses the content and parameter configuration of the automated test execution script. The task scheduler assigns the test task to the specified edge device node according to the preset execution strategy, triggers the execution process, and sends the script issuance instruction to the target edge device through the network communication protocol. After receiving the script file, the edge device performs local storage and permission configuration. When starting the vehicle system test, the edge device checks the connection status and power status of the vehicle device, and starts executing the operation instructions in the test script after confirming that the vehicle is in normal working condition. The test execution status feedback includes the script start time, current execution progress, operation success status and abnormal error information. The edge device sends the execution status information back to the Jenkins scheduling center through the real-time communication link. The scheduling center determines whether the test is proceeding normally based on the status feedback.

[0075] Multimodal data acquisition and processing synchronously activates various data acquisition devices based on the script startup signal in the test execution status feedback. When the industrial camera is started, the lens focus and exposure parameters are adjusted to ensure that the car screen is clear and visible. The image acquisition continuously captures the pixel data of the car screen at a fixed frame rate. Each frame of image data contains a timestamp, resolution information and RGB pixel value matrix. When the microphone is started, the audio gain and filter parameters are set, and the audio output signal of the car system is continuously collected. The audio data records the amplitude changes of the sound waveform at a fixed sampling frequency. Each sampling point contains a timestamp and amplitude value. The OBD interface is connected to the CAN bus network of the car and monitors the data frame transmission on the bus in real time. The CAN bus communication frame contains a frame identifier, data length, data content and check code. The original test data stream organizes the information of the three data sources into a continuous data sequence according to the acquisition time sequence.

[0076] The timestamp alignment process uniformly calibrates the clock differences of different data sources. Industrial cameras, microphones, and OBD interfaces each maintain an independent time base. The alignment algorithm identifies the synchronization event markers in each data source. The synchronization events include the execution time of specific operation instructions in the test script. The clock correction parameters are determined by calculating the time deviation of the synchronization events in different data sources. The corrected timestamps uniformly use the master clock of the edge device as the benchmark. The data format standardization process converts raw data of different formats into a unified data structure. Visual data is converted into a standard format containing timestamps, frame numbers, and pixel matrices. Audio data is converted into a standard format containing timestamps, sampling rates, and waveform arrays. Bus data is converted into a standard format containing timestamps, frame types, and data payloads. The synchronization integration process merges the three types of data into a single time series data stream according to a unified time base. Time-aligned multimodal data ensures that information of different modes accurately corresponds in the time dimension.

[0077] The feature extraction algorithm extracts key information from visual data, audio data, and bus data. Visual feature extraction detects changing areas of screen content through inter-frame difference calculations. The change detection algorithm calculates the difference in pixel values ​​between adjacent frames. Pixels with differences exceeding a threshold are marked as changing areas. The area and location information of the changing areas constitute the visual feature vector. Audio feature extraction uses spectral analysis to identify the frequency components and energy distribution of audio signals. Fast Fourier transform converts time-domain audio signals into frequency-domain representations. Spectral features include dominant frequency, harmonic components, and energy density. Bus feature extraction identifies control signals and status information in CAN frames through protocol parsing. The parsing process converts raw frame data into specific vehicle status parameters based on a predefined data dictionary. The anomaly detection algorithm establishes a baseline model of normal test behavior. The baseline model includes visual change patterns, audio output characteristics, and bus communication frequency under normal conditions. The real-time detection process compares the current features with the baseline model. Deviations exceeding a preset threshold are marked as abnormal events. Response delay identification is achieved by calculating the difference between the time it takes to send an operation command and the time it takes the interface to respond. The multi-dimensional test data set integrates all extracted feature information and detection results.

[0078] For example, in the vehicle navigation function test, the Jenkins scheduler sends a test script containing the operations of clicking the navigation icon and entering the destination to the edge device. After receiving the script, the edge device returns the script startup confirmation status to Jenkins, and at the same time activates the industrial camera to start recording the vehicle screen, the microphone starts collecting the voice prompt audio of the vehicle, and the OBD interface starts monitoring the communication data between the vehicle and the vehicle control module. When the test script executes the operation of clicking the navigation icon, the industrial camera captures the screen change from the main interface to the navigation interface, the microphone collects the prompt sound effect when the navigation is started, and the OBD interface monitors the sound sent by the vehicle to the GPS module. Positioning request instruction, the timestamp alignment algorithm recognizes the click operation as a synchronization event, calibrates the timestamps of the three data sources at the time of the event, and standardizes the data format to convert the camera's image frame, audio waveform data and OBD's CAN frame into a unified format. The feature extraction algorithm identifies the interface switching features from the image changes, extracts the spectral features of the prompt sound from the audio, and parses the GPS positioning request signal from the CAN frame. The anomaly detection algorithm compares the interface switching time with the expected response time to identify whether there is a response delay anomaly, and finally generates a multi-dimensional test data set containing interface response features, audio output features and bus communication features.

[0079] In a specific embodiment, the step of performing comprehensive analysis and processing of the time-series aligned multimodal data using feature extraction and anomaly detection algorithms may specifically include the following steps:

[0080] Perform frame-by-frame image difference calculation on the visual data in the time-aligned multimodal data, detect the vehicle interface response time through the pixel change rate, and obtain the interface response delay data;

[0081] Perform spectrum analysis and audio comparison on the audio data in the time-aligned multimodal data, extract audio features through short-time Fourier transform, calculate similarity, and obtain audio output quality assessment results.

[0082] Perform frame parsing and status monitoring on the CAN bus data in the time-aligned multimodal data, identify body control signals and fault codes, and obtain bus communication status data;

[0083] The interface response delay data, audio output quality assessment results and bus communication status data are fused and processed, and the comprehensive test indicators are calculated through a weighted algorithm to obtain a multi-dimensional test data set.

[0084] Specifically, frame-by-frame image difference calculation and processing performs continuous inter-frame comparative analysis on the visual data in the time-aligned multimodal data. The difference algorithm takes the corresponding pixel points of two adjacent frames of images and performs numerical subtraction operations. The difference is calculated for each pixel point's three RGB color channels respectively. The absolute value of the difference represents the degree of change of the pixel point between the two frames. The pixel change rate detection quantifies the intensity of interface change by counting the proportion of all pixel points whose change degree exceeds the preset threshold to the total number of pixels. The vehicle interface response time detection identifies the time difference between the execution time of the operation instruction and the time when the interface begins to change significantly. The operation instruction execution time is obtained from the timestamp record of the test script. The start time of the interface change is determined by detecting the frame timestamp when the pixel change rate first exceeds the response threshold. The difference between the two timestamps constitutes the interface response delay data, which reflects the response speed of the vehicle interface to user operations and the system performance status.

[0085] Spectral analysis processing performs frequency domain transformation and feature extraction on the audio data in the time-aligned multimodal data. The short-time Fourier transform divides the continuous time-domain audio signal into multiple time windows. The audio data in each time window is converted into frequency domain representation through fast Fourier transform. The frequency domain data contains the amplitude and phase information of different frequency components. Audio feature extraction calculates characteristic parameters such as the main frequency, harmonic structure, spectral center of gravity and frequency band energy distribution from the spectrum data. The audio comparison processing calculates the similarity between the audio features collected in real time and the features of the expected audio template. The similarity algorithm quantifies the degree of matching of the audio content by calculating the cosine distance or Euclidean distance between the two feature vectors. The smaller the distance value, the higher the audio similarity. The audio output quality evaluation result comprehensively considers indicators such as spectrum integrity, signal-to-noise ratio and distortion level to evaluate the accuracy and sound quality performance of the vehicle audio output.

[0086] The CAN bus data frame parsing process performs protocol decoding and information extraction on the bus data in the time-aligned multimodal data. The frame parsing parses the structure of each data frame according to the CAN protocol standard, including the frame start bit, arbitration field, control field, data field and check field. The identifier in the arbitration field is used to distinguish different information types and sending nodes. The data field contains actual control instructions or status information. The status monitoring process converts the original hexadecimal data into specific body control signals according to the predefined data dictionary. The body control signals include vehicle operating parameters such as engine speed, vehicle speed, steering angle, brake status, and light status. Fault code recognition detects abnormal conditions in the communication process between the vehicle and the computer by detecting specific diagnostic identifiers and error codes. The bus communication status data integrates all parsed control signals, status parameters and fault information to reflect the communication quality and functional normality between the vehicle and the computer.

[0087] The fusion calculation processing comprehensively analyzes and calculates indicators of the interface response delay data, audio output quality evaluation results and bus communication status data. The weighted algorithm assigns weight coefficients according to the importance of different test dimensions. The interface response performance weight reflects the importance of user interaction experience, the audio quality weight reflects the criticality of multimedia functions, and the bus communication weight represents the reliability requirements of the integration of the vehicle computer and the vehicle. The comprehensive test index is calculated by multiplying each data by the corresponding weight and then summing them up. The interface response delay data in the calculation formula needs to be normalized and converted into a response performance score. The shorter the response time, the higher the corresponding score. The audio quality evaluation result is already a standardized quality score. The bus communication status data is converted into a communication reliability score by counting the proportion of normal communication frames. The multidimensional test data set includes the original test data, intermediate processing results and the final comprehensive index to obtain a vehicle computer test evaluation system.

[0088] In a specific embodiment, step S5 includes:

[0089] Perform statistical analysis on the interface change data in the multi-dimensional test data set, calculate the number of successful executions and failures of test cases, and obtain the test pass rate statistics;

[0090] Perform time series analysis based on the response delay data in the multi-dimensional test data set, and obtain the vehicle system response time index through time difference calculation and average value statistics;

[0091] The test pass rate statistics and response time indicators are input into a lightweight text generation model based on the BERT architecture for report generation. The test result description text is generated according to the preset template structure to obtain the test analysis content.

[0092] The test analysis content is formatted and charted, and a pass rate pie chart, response time trend chart, and defect distribution statistics table are inserted to obtain a vehicle computer function verification report.

[0093] Specifically, statistical analysis and processing classify and count the interface change data in the multidimensional test data set and calculate the success rate. The interface change data contains the screen response status information during the execution of each test case. The statistics of the number of successful executions of the test case are determined by checking the status identifiers in the interface change data. When the interface change conforms to the expected pattern and is completed within a reasonable time range, it is marked as a successful execution. The success identifier includes conditions such as correct interface switching, normal display of target elements, and timely response to operation feedback. The failure count statistics identify abnormal interface changes, response timeouts, or incorrect interface displays. Failure identifiers include abnormal states such as interface freezing, error page jumps, and element loading failures. The calculation process traverses the execution records of all test cases, accumulates the number of success identifiers and failure identifiers, and the test pass rate statistics are calculated by dividing the number of successful executions by the total number of executions. This ratio reflects the stability and reliability level of the vehicle-computer interface function.

[0094] Time series analysis and processing extracts time features and performs statistical calculations on response delay data in a multidimensional test data set. The response delay data records the time interval from the sending of each operation instruction to the completion of the interface response. The time difference calculation extracts the operation start timestamp and the response completion timestamp and subtracts them to obtain the response delay of a single operation. The response delay data of all test operations form a time series. The average value statistics add all the values ​​in the time series and divide them by the total number of data points to calculate the average response time. The maximum and minimum value statistics identify the extreme values ​​of the response time. The standard deviation calculation reflects the fluctuation degree and stability of the response time. The vehicle system response time index integrates the statistical results of the average response time, the response time fluctuation range and the abnormal response events. The index data reflects the vehicle hardware performance and the degree of software optimization.

[0095] The lightweight text generation model receives test pass rate statistics and response time indicators as input data for report generation based on the BERT architecture. The encoder of the BERT model converts structured test data into semantic vector representation. The encoding process learns the association between test data through a multi-layer Transformer structure. The preset template structure defines the standard format and content framework of the report. The template contains fixed chapters such as test overview, execution results, performance analysis, and problem summary. The report generation process selects corresponding descriptive statements based on the numerical range and distribution characteristics of the input data. The text generation algorithm converts numerical test results into natural language descriptions. The generation process adjusts the emphasis of the language expression considering the severity and impact scope of the test results. The test analysis content includes the evaluation of the test pass rate, analysis of the response time performance, and detailed description of the problems found. The content generation follows the professional expression specifications and logical structure requirements of the technical report.

[0096] The formatting and typesetting processing performs layout design and adds visual elements to the test analysis content. The typesetting algorithm allocates page space and font style according to the length and importance of the content. The chart generation processing converts numerical test data into visual graphic elements. The pass rate pie chart shows the proportional relationship between successful and failed test cases. The pie chart drawing algorithm calculates the angle size of each sector area according to the pass rate value. The successful part is marked in green and the failed part is marked in red. The response time trend chart shows the change curve of the response time during the test. The trend chart is drawn with the test time as the horizontal axis and the response delay as the vertical axis. Connecting the data points to obtain a line graph. The defect distribution statistics table summarizes the various problems found during the test and their frequency of occurrence. The table contains column information such as defect type, number of occurrences, degree of impact and repair suggestions. The vehicle computer function verification report integrates text content and chart elements to obtain a complete technical document including cover, directory, main text and appendix.

[0097] The above describes the vehicle computer test method based on the small model in the embodiment of the present application. The following describes the vehicle computer test system based on the small model in the embodiment of the present application. Figure 2 In the embodiment of the present application, an embodiment of the vehicle computer test system based on a small model includes:

[0098] The extraction module is used to obtain the vehicle screen display image, extract the buttons, icons and text elements on the screen through a lightweight recognition model, and generate interface element recognition results including coordinate positions and element types;

[0099] An analysis module is used to analyze the functional association and operation sequence between elements based on the interface element recognition results, determine the touch click, slide switch and menu navigation operation types, and obtain the vehicle computer test operation sequence;

[0100] A generation module is used to generate ADB touch screen control instructions and robotic arm execution actions according to the operation type and coordinate parameters in the vehicle computer test operation sequence, and build an automated test execution script;

[0101] An operation module is used to run the automated test execution script to drive the vehicle system to execute test cases, and at the same time collect vehicle response data through industrial cameras, microphones and OBD interfaces to obtain a multi-dimensional test data set;

[0102] The calculation module is used to analyze the interface changes, audio output and bus communication status in the multi-dimensional test data set, calculate the test pass rate and response time indicators, and generate a vehicle function verification report.

[0103] above Figure 2The vehicle computer test system based on a small model in the embodiment of the present invention is described in detail from the perspective of modular functional entities. The vehicle computer test equipment based on a small model in the embodiment of the present invention is described in detail from the perspective of hardware processing.

[0104] Reference Figure 3 In an embodiment of the present invention, a vehicle computer test device based on a small model is also provided. The vehicle computer test device based on a small model can be a server, and its internal structure can be as follows: Figure 3 As shown. The vehicle-mounted test equipment based on the small model includes a processor, a memory, a display screen, an input device, a network interface and a database connected via a system bus. Among them, the computer-designed processor is used to provide computing and control capabilities. The memory of the vehicle-mounted test equipment based on the small model includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the vehicle-mounted test equipment based on the small model is used to store the corresponding data in this embodiment. The network interface of the vehicle-mounted test equipment based on the small model is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the above method is implemented.

[0105] Those skilled in the art will understand that Figure 3 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present invention, and does not constitute a limitation on the small-model-based vehicle-mounted test equipment to which the solution of the present invention is applied.

[0106] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions are executed on a computer, the computer executes the steps of the small-model-based vehicle-machine testing method.

[0107] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0108] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling a small-scale vehicle-mounted test device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0109] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A vehicle computer testing method based on a small model, characterized in that: The method comprises: Step S1: Obtain the vehicle screen display image, extract the buttons, icons and text elements on the screen through a lightweight recognition model, and generate an interface element recognition result including coordinate position and element type; Step S2: Analyzing the functional association and operation sequence between elements based on the interface element recognition results, determining the touch click, slide switch and menu navigation operation types, and obtaining the vehicle computer test operation sequence; Step S3: Generate ADB touch screen control instructions and robotic arm execution actions according to the operation type and coordinate parameters in the vehicle computer test operation sequence, and build an automated test execution script; Step S4: running the automated test execution script to drive the vehicle system to execute the test case, while collecting vehicle response data through the industrial camera, microphone and OBD interface to obtain a multi-dimensional test data set; Step S5: Analyze the interface changes, audio output and bus communication status in the multi-dimensional test data set, calculate the test pass rate and response time indicators, and generate a vehicle computer function verification report.

2. The vehicle computer testing method based on a small model according to claim 1 is characterized in that: The step S1 comprises: Capture the vehicle screen display image in real time through the HDMI interface of the edge device, convert the image data into standard RGB format and perform size normalization to obtain a standardized vehicle screen image; A lightweight convolutional neural network based on the MobileNetV3 architecture is used to extract features from the standardized vehicle screen image, and an image feature map is generated through a depthwise separable convolution and linear bottleneck structure. The image feature map is weighted using the channel attention mechanism, channel weights are calculated through global average pooling and a fully connected layer, and interface element features are output; The soft label knowledge of the teacher network is transferred to the student network through knowledge distillation technology, and the probability distribution is calculated using the softmax function to generate the element classification probability; The identified interface elements are post-processed to extract the bounding box coordinates, confidence score and category label of each element, and output structured interface element recognition result data.

3. The vehicle computer testing method based on a small model according to claim 1 is characterized in that: The step S2 includes: Input the coordinate position and element type in the interface element recognition result into a lightweight sequence modeling network based on the Transformer architecture for encoding processing to obtain a semantic feature vector of the interface element; Performing position encoding and temporal modeling on the semantic feature vectors of the interface elements, calculating the spatial association weights between elements through a multi-head attention mechanism, and obtaining an element association matrix; Performing grouping and clustering processing on the functional modules of the vehicle computer interface based on the element association matrix, identifying the menu navigation path and the interaction logic chain, and obtaining a functional operation path diagram; An operation sequence is generated based on the node connection relationship in the functional operation path diagram, and the click, slide, and long press operation types are matched and combined with the target coordinates to obtain the vehicle computer test operation sequence.

4. The vehicle computer testing method based on a small model according to claim 1 is characterized in that: The step S3 comprises: Parsing the operation type and coordinate parameters in the vehicle computer test operation sequence, matching the preset ADB command template according to the operation type to perform command splicing processing, and obtaining an ADB touch screen control command set; Calculating the joint angles and motion trajectories of the robotic arm based on the coordinate parameters in the vehicle-computer test operation sequence, performing spatial coordinate transformation processing through an inverse kinematics algorithm, and obtaining the robotic arm execution action instructions; The ADB touch screen control instruction set and the robotic arm execution action instruction are scripted according to the time sequence relationship, and waiting time and synchronization identifiers are inserted to obtain a time-sequential test instruction sequence; The sequential test instruction sequence is subjected to syntax checking and logic verification processing, and an exception handling and retry mechanism is added to obtain the automated test execution script.

5. The vehicle computer testing method based on a small model according to claim 1 is characterized in that: The step S4 comprises: The automated test execution script is triggered and executed through the Jenkins task scheduling system, test instructions are sent to the edge device, and the vehicle system test is started to obtain test execution status feedback; Based on the test execution status feedback, the industrial camera, microphone and OBD interface are synchronously started to perform multimodal data acquisition and processing, and the vehicle screen image, audio output and CAN bus communication frame are recorded in real time to obtain the original test data stream; Performing timestamp alignment and data format standardization on the original test data stream, synchronously integrating visual data, audio data, and bus data according to a unified time reference, and obtaining time-aligned multimodal data; The time-aligned multimodal data is comprehensively analyzed and processed through feature extraction and anomaly detection algorithms to identify response delays and functional abnormality events during the test process, thereby obtaining the multidimensional test data set.

6. The vehicle computer testing method based on a small model according to claim 5 is characterized in that: The multimodal data aligned with the time series is comprehensively analyzed and processed through feature extraction and anomaly detection algorithms to identify response delays and functional abnormalities during the test process, thereby obtaining the multidimensional test data set, including: Performing frame-by-frame image difference calculation processing on the visual data in the time-aligned multimodal data, detecting the vehicle interface response time by pixel change rate, and obtaining interface response delay data; Performing spectrum analysis and audio comparison processing on the audio data in the time-series aligned multimodal data, extracting audio features through short-time Fourier transform and calculating similarity to obtain an audio output quality assessment result; performing frame parsing and status monitoring processing on the CAN bus data in the time-aligned multimodal data, identifying body control signals and fault codes, and obtaining bus communication status data; The interface response delay data, audio output quality evaluation results and bus communication status data are fused and calculated, and comprehensive test indicators are calculated using a weighted algorithm to obtain the multi-dimensional test data set.

7. The vehicle computer testing method based on a small model according to claim 1 is characterized in that: The step S5 comprises: Performing statistical analysis on the interface change data in the multidimensional test data set, calculating the number of successful executions and failures of the test cases, and obtaining test pass rate statistics; Performing time series analysis on the response delay data in the multi-dimensional test data set, and obtaining a vehicle system response time index by calculating the time difference and averaging the average value; Input the test pass rate statistics and response time indicators into a lightweight text generation model based on the BERT architecture for report generation processing, generate test result description text according to a preset template structure, and obtain test analysis content; The test analysis content is formatted and charted, and a pass rate pie chart, a response time trend chart, and a defect distribution statistics table are inserted to obtain the vehicle computer function verification report.

8. A vehicle computer test system based on a small model, characterized in that: For implementing the vehicle computer testing method based on a small model according to any one of claims 1 to 7, the vehicle computer testing system based on a small model comprises: The extraction module is used to obtain the vehicle screen display image, extract the buttons, icons and text elements on the screen through a lightweight recognition model, and generate interface element recognition results including coordinate positions and element types; An analysis module is used to analyze the functional association and operation sequence between elements based on the interface element recognition results, determine the touch click, slide switch and menu navigation operation types, and obtain the vehicle computer test operation sequence; A generation module is used to generate ADB touch screen control instructions and robotic arm execution actions according to the operation type and coordinate parameters in the vehicle computer test operation sequence, and build an automated test execution script; An operation module is used to run the automated test execution script to drive the vehicle system to execute test cases, and at the same time collect vehicle response data through industrial cameras, microphones and OBD interfaces to obtain a multi-dimensional test data set; The calculation module is used to analyze the interface changes, audio output and bus communication status in the multi-dimensional test data set, calculate the test pass rate and response time indicators, and generate a vehicle function verification report.

9. A vehicle computer testing device based on a small model, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the computer program, the vehicle-computer testing method based on a small model according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the processor is enabled to execute the vehicle computer testing method based on a small model according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Intelligent vehicle machine mechanical arm test platform system

    CN113900960A

  • Distributed optical fiber sensing pipeline leakage identification method

    CN119123343A

  • Software automatic testing method and system based on visual attention mechanism

    CN119292953A

  • Automatic test script dynamic generation method and system based on multi-modal AI identification

    CN120011247A

Cited By

  • Intelligent detection method, system and equipment for core control board of distribution network hot-line work robot and storage medium

    CN120949754A

  • Power distribution network live working robot core control board intelligent detection method, system, equipment and storage medium

    CN120949754B

  • Method and device for identifying and positioning multi-modal large-model vehicle touch test keys

    CN121541817A

  • Autonomous railway fastener assembling and positioning method and system based on multi-modal visual fusion

    CN122089842A