A method and system for automatic testing of set-top boxes based on big data

By building a lightweight agent in the set-top box for log data acquisition and big data analysis, using K-means clustering and lightweight random forest model to generate dynamic test cases, solving the problems of insufficient coverage of test cases and static strategies in set-top box automation tests, improving the test coverage and accuracy, and achieving dynamic optimization of user behavior.

CN119854576BActive Publication Date: 2025-07-11海看网络科技(山东)股份有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510338277.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-11
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

The existing set-top box automated tests have problems such as insufficient coverage of test cases, static test strategies, low data utilization, and mismatch between test efficiency and accuracy. Especially when user behavior changes rapidly, the existing technology fails to effectively optimize test strategies and utilize big data analysis.

Method used

Log data collection is collected through the built-in lightweight agent of the device system, combined with big data analysis technology, using K-means clustering and lightweight random forests to build defect prediction models, generate core and boundary test cases, dynamically optimize test strategies, and improve test scenario coverage and accuracy.

Benefits of technology

The test cases are more in line with the user's real usage scenarios, and the test strategies can be dynamically optimized, which improves test coverage and accuracy, and makes full use of device log data for optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119854576B_ABST
    Figure CN119854576B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for automatic testing of set-top boxes based on big data, mainly relating to the technical field of automatic testing of set-top boxes. It includes: collecting logs through a lightweight agent built into the device system and extracting log data information; cleaning the extracted data information through filtering rules to form a data set; using K-means clustering to extract behavior patterns from the data set after data cleaning and constructing a defect prediction model; dividing the extracted log data information into a training set and a validation set and training using the defect prediction model; generating core test cases and boundary test cases. The beneficial effect of the present invention is that it can achieve that the test cases are more in line with the real usage scenarios of users while dynamically optimizing the test strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of set-top box automated testing, and specifically, to a method and system for set-top box automated testing based on big data. Background Art

[0002] Currently, set-top box automated testing mainly writes or records test cases through preset test scenarios, and based on predefined remote control key commands (such as "right arrow → down arrow → confirm") and UI screenshot comparison to cover fixed operation paths.

[0003] Using the above method for multiple devices can indeed improve the testing efficiency, but there are also the following problems:

[0004] 1. Insufficient test case coverage: Existing solutions rely on predefined test cases, which are usually based on standard scenarios or developers' assumptions and cannot fully reflect the usage habits of real users. For example, users may frequently use the live broadcast function during a specific time period (such as weekend evenings), but such scenarios are not covered by the test cases.

[0005] 2. Static test strategy: Once the test cases and strategies are recorded, subsequent adjustments are rather troublesome and cannot be dynamically optimized according to changes in user behavior, which will cause the test results to gradually deviate from the actual usage scenarios, especially in the context of rapidly changing user habits.

[0006] 3. Low data utilization rate: Although the test platform can collect set-top box operation data (such as logs, error reports), these data are only used to generate static reports and are not deeply analyzed and applied to optimize the test process, wasting potential value.

[0007] 4. Mismatch between test efficiency and accuracy: Existing technologies usually adopt a unified test strategy when testing a large number of devices, ignoring individual differences (such as network conditions and device model differences of users in different regions), resulting in inaccurate test results and low efficiency in some scenarios.

[0008] The main reasons for these problems are as follows:

[0009] 1. Lack of dynamic feedback mechanism: Existing systems do not effectively link actual user usage data with the test process, resulting in overly static test case design and inability to adapt to diverse user behaviors.

[0010] 2. Insufficient data analysis ability: Although existing test platforms have data collection functions, they lack big data processing and analysis capabilities and cannot extract meaningful data from massive data.

[0011] 3. Low degree of technology integration: The automation testing and big data analysis technologies are not deeply integrated in the existing solutions, which restricts the intelligence and adaptability of the testing process.

[0012] Therefore, there is an urgent need for a set-top box automation testing method and system based on big data to solve the above problems. Summary of the Invention

[0013] The purpose of the present invention is to provide a set-top box automation testing method and system based on big data, which can achieve dynamic optimization of the testing strategy while making the test cases more conform to the real usage scenarios of users.

[0014] To achieve the above object, the present invention is realized through the following technical solutions:

[0015] On the one hand, the present invention provides a set-top box automation testing method based on big data, including the following steps:

[0016] S1: Collect log data through a lightweight agent built into the device system;

[0017] S2: Clean the extracted data information through filtering rules to form a data set;

[0018] S3: Use K-means clustering to extract behavior patterns from the data set after data cleaning, and use a lightweight random forest to construct a defect prediction model;

[0019] S4: Divide the log data information extracted in step S1 into a training set and a validation set, and use the defect prediction model for training;

[0020] S5: Generate core test cases and boundary test cases.

[0021] Preferably, step S1 includes:

[0022] S11: Extract the key information of the log, and the key information includes: device id, time point, operation behavior;

[0023] S12: Set the data collection interval time and the number of abnormal retries;

[0024] S13: Execute instructions for data collection and data transmission alternately.

[0025] Preferably, step S2 includes:

[0026] S21: Remove records with missing or incorrect timestamp formats, delete duplicate data and invalid operation paths. If two records have the same time, keep the latest one;

[0027] S22: Set the encoding mapping table for the operation path.

[0028] S23: Extract time series features, calculate the average operation interval time for the last ten times and the cumulative usage duration of the day;

[0029] S24: Normalize values such as duration and network latency to ;

[0030] S25: Evaluate the health of the device through log data:

[0031] ;

[0032] Among them, is the health evaluation of the device, and a warning reminder is triggered when .

[0033] Preferably, for the dataset after performing data cleaning in step S3, using K-means clustering to extract behavior patterns includes:

[0034] Initialize K clustering centers, and calculate the Euclidean distance from each data point to the center: The distance formula is: ;

[0035] Assign the data points to the nearest center, update the center position to the mean within the cluster, and iterate until the center no longer changes or reaches the maximum number of iterations;

[0036] Output 5 types of behavior patterns;

[0037] Determine the optimal K value through the elbow method, based on the WSS curve.

[0038] Preferably, step S4 includes:

[0039] S41: Divide the log data information extracted in step S1 into a training set and a validation set in a ratio of 8:2 for training, and determine the optimal tree depth and number through grid search;

[0040] S42: Add the daily test result data to the training set, and perform full-scale retraining when the accuracy rate drops by more than 2%;

[0041] S43: Perform time series anomaly detection, calculate the daily network latency mean and standard deviation , and calculate the value for each record: , if , then mark it as an anomaly;

[0042] S44: Construct a priority algorithm:

[0043] ;

[0044] S45: Build a test task generation rule library and execute it in descending order of priority.

[0045] Preferably, the step S5 is specifically: based on the behavior pattern, mapping table, and priority sorting, generate core test cases and boundary test cases respectively. The core test cases are based on high-frequency behavior patterns, and the boundary test cases are based on abnormal patterns.

[0046] Preferably, it also includes a detection process for test cases, including:

[0047] Through the API interface, check the syntax and executability of the test cases. If it fails, record the log and retry the generation;

[0048] Receive the successfully generated test cases, and according to parameters such as device model, performance, and preset rules, allocate and combine the test cases and devices;

[0049] Each proxy node conducts concurrent testing on a device, split the received test cases, split the test cases into test steps, identify the device model corresponding to the case, and perform specified distribution;

[0050] During the test execution process, the proxy node continuously feeds back the running logs and abnormal conditions of the device to the automated test platform, and automatically generates a visual report and result analysis after the test is completed.

[0051] On the other hand, a test system based on the above-mentioned big data-based set-top box automated test method is provided, including:

[0052] A data collection module for: collecting log data through a lightweight agent built into the device system;

[0053] A data processing module for: cleaning the extracted data information through filtering rules to form a data set;

[0054] A model construction module for: using K-means clustering to extract behavior patterns from the data set after data cleaning, and using a lightweight random forest to build a defect prediction model;

[0055] A model training module for: dividing the extracted log data information into a training set and a validation set, and training using the defect prediction model;

[0056] A data output module for: generating core test cases and boundary test cases.

[0057] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0058] 1. The test cases are more in line with the real user usage scenarios: The present invention utilizes big data to analyze the operation logs of the device, obtains high-frequency and special operation behaviors, and combines with an automated test platform to generate test cases with different priorities, improving the coverage rate of the test scenarios and the accuracy of the tests.

[0059] 2. The test strategy is dynamically optimized: The loop mechanism of the feedback optimization unit of the present invention enables the test strategy to be continuously adjusted dynamically according to the test results, making it more accurate and reasonable.

[0060] 3. The full utilization of log data: The present invention combines data collection and big data analysis, making the operation logs of the device more valuable, rather than just being used for the report display of conventional automated tests. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 is the flowchart of the method of the present invention;

[0062] Figure 2 is the flowchart of data preprocessing of the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0063] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by this application.

[0064] In the present invention, terms such as "upper", "lower", "left", "right", "front", "rear", "vertical", "horizontal", "side", "bottom", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only relationship terms determined for the convenience of describing the structural relationship of each component or element of the present invention, rather than specifically referring to any component or element in the present invention, and should not be construed as a limitation to the present invention.

[0065] Embodiment:

[0066] As Figure 1 shown, this embodiment provides an automated test method for a set-top box based on big data, including the following steps:

[0067] S1: Collect log data through a lightweight agent built into the device system;

[0068] S2: Clean the extracted data information through filtering rules to form a data set;

[0069] S3: Use K-means clustering to extract behavior patterns from the data set after data cleaning, and use a lightweight random forest to construct a defect prediction model;

[0070] S4: Divide the log data information extracted in step S1 into a training set and a validation set, and use the defect prediction model for training;

[0071] S5: Generate core test cases and boundary test cases.

[0072] In step S1, lightweight agents built into the device system are used for log data collection, including:

[0073] S11: Extract the key information of the log. The key information includes: device id, time point, operation behavior, etc. It can also include metadata such as device model, firmware version, and geographical location to support more fine-grained analysis. Wrap the log into a preset format, encrypt sensitive information using AES-256, and use SHA-256 verification to prevent data tampering. In this embodiment, a log example is as follows:

[0074] {

[0075] "device_id": "STB001",

[0076] "timestamp": "2025-02-10 10:00:00",

[0077] "channel_switch": 5,

[0078] "app_usage": {"vod": 1200, "live": 300},

[0079] "network_delay": 50,

[0080] "errors": ["buffer_timeout"]

[0081] };

[0082] S12: Set the data collection interval time and the number of abnormal retries. Specifically: the data collection interval is set to 5 minutes, and abnormal retries are 3 times. After the collection is completed, ensure that the data packet is MB. During peak hours (such as at night), the data volume may increase, and the agent automatically adjusts the frequency to 3 minutes to avoid network congestion caused by overly large data packets;

[0083] S13: Execute instructions for alternating data collection and data sending. Specifically: within the task interval time after data collection is completed, send the data to the big data analysis platform. The connection timeout time with the platform is 10s, the retry interval is 5 minutes, and the maximum number of retries is 3 times. Cache locally after failure. The maximum cache is 10MB, and resend after the network is restored.

[0084] Step S2 is specifically as follows:

[0085] The big data platform cleans the massive data through filtering rules, removes records with missing timestamps or incorrect formats, deletes duplicate data (based on device_id and timestamp) and invalid operation paths. If two records have the same time, the latest one is retained;

[0086] Set the operation path coding mapping table, such as "power on" = 1, "on demand" = 2, "movie" = 3, etc.;

[0087] Extract time series features, calculate the average operation interval time of the last 10 times and the cumulative usage duration of the day, etc.;

[0088] Normalize values such as duration and network latency to [0, 1]: formula: : The network latency range is [10, 200], and 50 is normalized to (50 - 10) / (200 - 10) = 0.21. Extract derivative features, such as the total daily usage duration and average network latency;

[0089] Evaluate the health of the device through data, and the calculation rules are as follows:

[0090]

[0091] When HealthScore < 0.6, a warning reminder is triggered.

[0092] Step S3 is specifically as follows:

[0093] Use K-means clustering on the cleaned data set to extract behavior patterns. The features include path, switching frequency, duration, network latency, etc. The main process is as follows:

[0094] 1) Initialize K clustering centers (K = 5, randomly selected), and calculate the Euclidean distance from each data point to the center: distance formula: ;

[0095] 2) Assign the data points to the nearest center, update the center position to the mean within the cluster, and iterate until the center no longer changes or reaches the maximum number of iterations (100 times);

[0096] 3) Output behavior patterns, such as: Cluster 1: Night live users (high live duration), Cluster 2: Daytime on-demand users (high on-demand duration), high-frequency switching users, low-latency users, peak period users, users with poor performance, etc. Support dynamic adjustment of the K value, and determine the optimal number of clusters through the Elbow Method;

[0097] 4) Based on the within-cluster sum of squares (WSS) curve, the optimal K value is determined by the elbow method (ElbowMethod);

[0098] Build a defect prediction model using lightweight random forest, number of decision trees = 50, maximum depth = 5. Feature vector: feature_vector = [

[0099] path_encoded, # Encoded operation path (such as [1,2,3])

[0100] health_score,# Health score (0.72)

[0101] network_type,# Network type (0=5G, 1=broadband)

[0102] time_since_last_reboot# Time since last reboot (8 hours)

[0103] ].

[0104] Step S4 is specifically:

[0105] The log data is trained with a ratio of 80% for training set and 20% for validation set, and grid search is used to determine the optimal tree depth and number. Output example: {"device_id": "STB-001", "P_defect": 0.85, "risk_type": "memory overflow"};

[0106] Add daily test result data to the training set, and retrain the entire set when the accuracy drops by more than 2%;

[0107] Perform time series anomaly detection (Z-score) and calculate the daily average network latency and standard deviation , calculate the Z value for each record: ,like , marked as abnormal (such as sudden increase in network delay), example: μ=50ms, σ=10ms, x=85ms, Z=(85-50) / 10=3.5, marked as abnormal;

[0108] Construct a priority algorithm, the calculation formula is:

[0109] ;

[0110] Represents the probability of defect occurrence, indicating the likelihood of a defect existing in the current test objective. The input feature vectors (operation path, health score, network type, etc.) are fed into the random forest model, and a probability value (ranging from 0 to 1) is output. For example: Indicates an 85% probability of triggering a defect; Represents the operation frequency, which is the number of times a certain test objective is triggered within a statistical period. For example, the path "channel switching" is triggered 50 times within 10 minutes; Represents the maximum frequency, that is, the highest number of trigger times among all test objectives within the same statistical period. The frequencies of all operation paths or devices within the current period are sorted, and the maximum value is taken; 0.7 and 0.3 are weight coefficients used to balance the defect risk and operation frequency;

[0111] That is: for a certain path where OperationFrequency = 150 times / hour (MaxFrequency = 200) and P_defect = 0.85, then Priority = 0.7×0.85 + 0.3×(150 / 200) = 0.805;

[0112] Construct a test task generation rule library and execute it from high to low priority:

[0113] {

[0114] "task_type": "stress test",

[0115] "trigger_condition": "HealthScore<0.6&&P_defect>0.8",

[0116] "parameters": {

[0117] "concurrency": 10,

[0118] "resolution": "3840x2160",

[0119] "duration": "300s"

[0120] }

[0121] }

[0122] Step 5, specifically:

[0123] Based on the calculated behavior pattern + rule mapping + priority sorting, generate core test cases and boundary test cases respectively. Core cases: based on high-frequency behavior patterns (such as frequent channel switching); Boundary cases: based on abnormal patterns (such as network interruption);

[0124] The process is as follows:

[0125] 1) Map the behavior patterns to the test objectives. For example: "Live streaming users at night" → Test the smoothness of live streaming, "Abnormal network latency" → Test the recovery from network disconnection;

[0126] 2) Calculate the priority:

[0127] Priority = Usage frequency × Abnormality weight (the default abnormality weight is 1, and the abnormality scenario weight is 2);

[0128] Example: The usage frequency of live streaming is 0.8, the abnormality weight is 1, and the priority = 0.8.

[0129] In this embodiment, the syntax and executability of test cases can also be checked through the API interface, supporting JavaScript, Python, etc. If it fails, record the log and retry to generate;

[0130] The automated testing platform receives the successfully generated test cases and allocates and combines the test cases and devices according to parameters such as device model, performance, and preset rules;

[0131] The test cases are distributed to the corresponding proxy nodes of the set-top box and wait for the proxy nodes to execute the tests;

[0132] Each proxy node concurrently tests less than 10 devices, splits the received test cases, splits the test cases into test steps, identifies the device model corresponding to the case, and performs specified distribution;

[0133] During the test execution process, the proxy node continuously feeds back the running logs and abnormal situations of the devices to the automated testing platform, and automatically generates a visual report and result analysis after the test is completed;

[0134] Regularly analyze the test results, calculate the failure rate = the number of failed cases / the total number of cases, and identify the scenarios with high failure rates (such as the proportion of cases with a response time > 1 second);

[0135] Combined with the analysis results of the previous step, adjust the weights, increase the clustering weight of the failure scenarios, and the new weight = the original weight × (1 + failure rate). Example: The original weight is 1, the failure rate is 0.2, and the new weight = 1.2;

[0136] Rerun K-means to generate new cases, and the weights can also be adjusted using reinforcement learning to predict the future optimization direction based on historical failure rates;

[0137] The test strategy can be optimized in combination with user feedback. For example, if it is found through customer service records that users feedback that the live streaming at night is stuck, the weight of the night live streaming test can be correspondingly increased.

[0138] The above is a specific description of the preferred embodiment of the present invention, but the present invention is not limited to the described embodiment. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.

Claims

1. An automated testing method for a set-top box based on big data, characterized in that, It includes the following steps: S1: Collect log data through the lightweight agent built into the device system; S2: Clean the extracted data information through filtering rules to form a data set; S3: Use K-means clustering to extract behavior patterns from the data set after data cleaning, and use a lightweight random forest to build a defect prediction model; S4: Divide the log data information extracted in step S1 into a training set and a validation set, and use the defect prediction model for training, including the following steps: S41: Divide the log data information extracted in step S1 into a training set and a validation set according to a ratio of 8:2 for training, and determine the best tree depth and number through grid search; S42: Add the daily test result data to the training set, and perform full-scale retraining when the accuracy drops by more than 2%; S43: Perform time series anomaly detection and calculate the daily network delay mean and standard deviation , and calculate for each record value: , if , then mark it as an anomaly; S44: Construct a priority algorithm: ; Among them, represents the defect occurrence probability, indicating the possibility that the current test target has a defect; represents the operation frequency, which is the number of times a certain test target is triggered within the statistical period; represents the maximum frequency, that is, the highest number of trigger times among all test targets within the same statistical period. 0.7 and 0.3 are weight coefficients used to balance the defect risk and the operation frequency; S45: Build a test task generation rule library and execute it from high to low priority; S5: Generate core test cases and boundary test cases.

2. The automated testing method for a set-top box based on big data according to claim 1, characterized in that The step S1 includes: S11: Extract the key information of the log, and the key information includes: device id, time point, operation behavior; S12: Set the data collection interval time and the number of abnormal retries; S13: Execute instructions for data collection and data sending alternately.

3. The automated testing method for a set-top box based on big data according to claim 1, wherein, The step S2 includes: S21: Remove records with missing or incorrectly formatted timestamps, delete duplicate data and invalid operation paths. If two records have the same time, keep the latest one; S22: Set the encoding mapping table for the operation path; S23: Extract time series features, calculate the average operation interval time of the last ten times and the cumulative usage duration of the day; S24: Normalize the duration and network latency values to ; S25: Evaluate the health of the device through log data: ; Among them, is the health assessment of the device, and a warning reminder is triggered when occurs.

4. The automated testing method for a set-top box based on big data according to claim 1, wherein The use of K-means clustering in step S3 to extract behavior patterns from the data set after data cleaning includes: Initialize K cluster centers and calculate the Euclidean distance from each data point to the centers: The distance formula is: ; Assign data points to the nearest center, update the center position to the mean within the cluster, and iterate until the center no longer changes or reaches the maximum number of iterations; According to the clustering results of the data set, output behavior patterns, and the behavior patterns include but are not limited to: night live users, daytime on-demand users, high-frequency switching users, low-latency users, peak users, users with poor performance; Based on the WSS curve, determine the optimal K value through the elbow method.

5. The automated testing method for a set-top box based on big data according to claim 3, characterized in that, The step S5 is specifically: Generate core test cases and boundary test cases based on behavior patterns, mapping tables, and priority sorting. The core test cases are based on high-frequency behavior patterns, and the boundary test cases are based on abnormal patterns.

6. The automated testing method for a set-top box based on big data according to claim 1, characterized in that It also includes a detection process for test cases, including: Check the syntax and executability of test cases through the API interface. If it fails, record the log and retry generation; Receive the successfully generated test cases, and allocate combinations of test cases and devices according to device models, performance, and preset rule parameters; Concurrent testing for each proxy node Split the received test cases for devices, split the test cases into test steps, identify the device model corresponding to the test case, and perform specified distribution; During the test execution process, the agent node continuously feeds back the running logs and abnormal situations of the device to the automated test platform, and automatically generates a visual report and result analysis after the test is completed.

7. A test system for a set-top box automated test method based on big data as described in claim 1, characterized in that, It includes: A data collection module for: Collecting log data through the lightweight agent built into the device system; A data processing module, which is used for: cleaning the extracted data information through filtering rules to form a data set; A model construction module, which is used for: extracting behavior patterns from the data set after data cleaning using K-means clustering, and constructing a defect prediction model using a lightweight random forest; A model training module, which is used for: dividing the extracted log data information into a training set and a validation set, and training using the defect prediction model; A data output module, which is used for: generating core test cases and boundary test cases.

Citation Information

Patent Citations

  • Code testing method and device, electronic equipment and storage medium

    CN118132444A

  • Automatic test method based on behavior-driven test

    CN118860894A