Set top box fault prediction method and device based on log data, medium and equipment
By analyzing the set-top box log data and training the fault prediction model, the traditional problem of inefficient troubleshooting is solved, and accurate prediction and efficient maintenance of potential faults are achieved.
Patent Information
- Application Number
- CN202510273747.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-10
AI Technical Summary
Traditional set-top box troubleshooting relies on user reports or on-site inspections, is inefficient and cannot prevent potential failures in advance.
By analyzing the log data of the set-top box, a fault prediction model is constructed and trained, and a smart sampler and a multimodal feature fusion network are used to predict.
Accurate prediction of potential failures of set-top boxes is achieved, maintenance efficiency and service quality are improved, the possibility of users encountering problems, and the reliability of the system is improved.
Smart Images

Figure CN120128696A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of data processing, and particularly relates to a set-top box fault prediction method, device, medium, and equipment based on log data. Background Art
[0002] With the development of digital TV and the Internet, as an important part of the home entertainment center, the stability and reliability of the set-top box have become particularly important. However, due to reasons such as hardware aging, software vulnerabilities, and changes in the network environment, various problems may occur in the set-top box, thus affecting the user experience. Traditional set-top box fault troubleshooting relies on user reports or on-site inspections by technicians. This method is not only inefficient but also often can only handle problems after they occur and cannot prevent them in advance.
[0003] Therefore, in order to improve the maintenance efficiency and service quality of the set-top box, a method that can automatically analyze the set-top box log data and predict potential faults is needed. Summary of the Invention
[0004] Aiming at the deficiencies in the prior art, the main purpose of this application is to provide a set-top box fault prediction method, device, medium, and equipment based on log data. By analyzing the log data, this application can predict whether there are potential faults in the set-top box.
[0005] To achieve the above objectives, this application provides the following technical solutions:
[0006] A set-top box fault prediction method based on log data, the method includes: obtaining set-top box log data within a set period, where the set-top box log data includes system error and warning information, application crash reports, network connection status changes, hardware health indicators, user behavior patterns, and performance statistics; preprocessing the log data; constructing a set-top box fault prediction model and training the model; the set-top box fault prediction model introduces an intelligent sampler to perform intelligent sampling on the preprocessed log data to reduce the data volume; inputting the preprocessed log data into the trained set-top box fault prediction model to predict the faults of the set-top box.
[0007] Optionally, preprocessing the log data includes: cleaning the log data; formatting the cleaned log data; performing feature engineering on the formatted log data.
[0008] Optionally, the set-top box fault prediction model includes: a backbone network, a feature fusion network, and an output network, where the backbone network is used to extract features from the preprocessed log data; the feature fusion network is used to fuse the features extracted by the backbone network to obtain comprehensive features; the output network is used to receive the comprehensive features and output the fault prediction result of the set-top box.
[0009] Optionally, the backbone network includes, connected in sequence: an input layer, an intelligent sampler, a hybrid embedding layer, an adaptive sampling layer, and a dynamic multi-resolution feature aggregation network.
[0010] Optionally, the feature fusion network includes, connected in sequence: a multi-modal input layer, a modality-specific feature extraction layer, a cross-modal interaction layer, and a global context awareness module.
[0011] Optionally, the output network includes, connected in sequence: a fully connected layer, a Dropout layer, a batch normalization layer, and an activation function.
[0012] Optionally, the set-top box fault prediction model is trained through the following steps: obtaining historical fault data of the set-top box and preprocessing it, dividing the preprocessed historical fault data into a training set and a validation set according to a ratio; setting training parameters, using the training set to train the model, and during the model training process, when the number of model iterations meets a preset value, the model training is completed; using the validation set to validate the trained model, and during the validation process, if the mean absolute error, mean square error, and root mean square error, which are used as model performance evaluation indicators, are all less than a threshold, the model validation passes; otherwise, adjust the training parameters or expand the training set samples to retrain the model until the validation passes.
[0013] This application also provides a set-top box fault prediction device based on log data. The device includes: an acquisition module, used to acquire the set-top box log data within a set period, where the set-top box log data includes system error and warning information, application crash reports, network connection status changes, hardware health indicators, user behavior patterns, and performance statistics; a preprocessing module, used to preprocess the log data; a model construction and training module, used to construct a set-top box fault prediction model and train the model; a prediction model, used to input the preprocessed log data into the trained set-top box fault prediction model to predict the faults of the set-top box.
[0014] This application also provides a storage medium, which includes instructions that, when running on a computer, cause the computer to execute the method described in any of the previous items.
[0015] The present application also provides an electronic device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. Wherein, when the processor executes the program, the method described in any of the previous paragraphs is implemented.
[0016] The present application can bring the following beneficial effects:
[0017] By constructing and training a set-top box fault prediction model, the present application can accurately predict potential faults of the set-top box, thereby improving the maintenance efficiency and service quality of the set-top box, reducing the possibility of users encountering problems. Moreover, compared with the traditional method that relies on user reports or on-site inspections, the present application has a high degree of automation, can respond and prevent faults more quickly, and can improve user experience and system reliability. Description of the Drawings
[0018] Figure 1 is a schematic flowchart of a method for predicting set-top box faults based on log data provided by an embodiment of the present application;
[0019] Figure 2 is a schematic structural diagram of a set-top box fault prediction model provided by another embodiment of the present application;
[0020] Figure 3 is a schematic structural diagram of a device for predicting set-top box faults based on log data provided by another embodiment of the present application;
[0021] Figure 4 is a schematic structural diagram of a storage medium provided by another embodiment of the present application;
[0022] Figure 5 is a schematic structural diagram of an electronic device provided by another embodiment of the present application. Detailed Embodiments
[0023] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0024] It should be noted that all directional indications (such as up, down, left, right, front, back...) in the embodiments of the present application are only used to explain the relative position relationship and movement conditions between components in a specific posture (as shown in the drawings). If the specific posture changes, the directional indications will also change accordingly.
[0025] In this application, unless otherwise clearly specified or limited, terms such as "connection" and "fixation" shall be understood in a broad sense. For example, "fixation" can be a fixed connection, a detachable connection, or integrated; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components or the interaction relationship between two components, unless otherwise clearly limited. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0026] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of this application, the descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" can explicitly or implicitly include at least one of such features. In addition, the meaning of "and / or" appearing throughout the text includes three parallel scenarios. Taking "A and / or B" as an example, it includes scenario A, scenario B, or the scenario where both A and B are satisfied simultaneously. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by this application.
[0027] Figure 1 It is a schematic flowchart of a method for predicting set-top box failures based on log data provided by an exemplary embodiment of this application, as Figure 1 shown, and the prediction method includes the following steps:
[0028] S100: Obtain the set-top box log data within a set period. The log data includes system errors and warning messages (for example, including the timestamp: the time when each error or warning occurred; error code: the specific error code or descriptive text generated by the system), application crash reports (for example, including the application name and version: record which application crashed and its version number; crash time and frequency: the time when each crash occurred and the number of times the application crashed within a specific period), network connection status changes (for example, including the connection / disconnection time: the exact time when the network connection was established or disconnected; reason for connection failure: if the connection attempt failed, record the reason for failure, such as DNS resolution failure, authentication failure, etc.), hardware health metrics (for example, including the temperature reading: the temperature change trend of the CPU / GPU or other key components; disk SMART status: the Self-Monitoring, Analysis and Reporting Technology status of the hard disk), user behavior patterns (for example, including the remote control operation frequency: the operation frequency of common functions such as channel switching and volume adjustment, an abnormally high frequency may indicate that the user has encountered a problem; playback interruption situation: any interruption events during video playback, including loading delays, frequent buffering, etc.) and performance statistics (for example, including the CPU usage peak: record the time point when the CPU utilization suddenly soars and its duration; memory occupancy peak: record the situation when the memory usage reaches the peak).
[0029] S200: Preprocess the log data;
[0030] S300: Build a set-top box fault prediction model and train the model;
[0031] S400: Input the preprocessed log data into the trained set-top box fault prediction model to predict the faults of the set-top box.
[0032] By building and training a set-top box fault prediction model, this application can accurately predict potential faults of the set-top box based on a neural network, thereby improving the maintenance efficiency and service quality of the set-top box and reducing the possibility of users encountering problems.
[0033] In another exemplary embodiment, in step S200, the preprocessing of the log data includes the following steps:
[0034] S201: Clean the log data;
[0035] In this step, cleaning the log data includes filtering out useless logs, removing duplicates, and filling in missing values. Among them, filtering out useless logs means removing those log entries that do not contain useful information, such as advertisements, promotional information, or other non-critical events. Removing duplicates means identifying and deleting exactly the same log records to avoid bias in subsequent analysis. Filling in missing values means that for cases where some fields are missing, the mean / median filling, forward fill, backward fill, or more complex prediction models can be used to fill these gaps.
[0036] S202: Format the cleaned log data;
[0037] In this step, for log content containing a large amount of free text, regular expressions, natural language processing tools, or dedicated log parsing libraries can be used to decompose it into a structured key-value pair form. When it comes to hardware health metrics (such as temperature, voltage, etc.), it is necessary to unify the measurement units to ensure that all measurement units are consistent, for example, all converted to degrees Celsius or volts. In addition, for some particularly large values (such as disk capacity), it can be converted into an easily understandable decimal form through an appropriate scaling factor. Further, categorical features (such as error codes, application names, etc.) also need to be converted into binary vector representations, so that each possible value corresponds to an independent dimension.
[0038] S203: Perform feature engineering on the formatted log data.
[0039] In this step, first, it is necessary to extract features from the formatted log data, specifically including aggregating statistical data and time window analysis. Among them, aggregating statistical data means calculating summary statistics of various performance metrics, such as mean, maximum, minimum, standard deviation, etc., to capture the overall behavior trend of the system. Time window analysis means calculating the activity frequency, duration, and change rate within a specific time period based on a fixed time interval (such as 5 minutes, 1 hour, one day) to reveal potential periodic and sudden patterns. Secondly, it is necessary to reduce the dimension of the extracted features, specifically through principal component analysis (PCA) to retain as much original information as possible while reducing the number of features.
[0040] In another exemplary embodiment, in step S300, as Figure 2 shown, the set-top box fault prediction model includes: a backbone network, a feature fusion network, and an output network. Among them, the backbone network is used to extract features from the preprocessed log data; the feature fusion network is used to fuse the features extracted by the backbone network; the output network is used to receive the comprehensive features processed by the feature fusion network and output the fault prediction result of the set-top box.
[0041] In this embodiment, the backbone network includes an input layer, an intelligent sampler, a hybrid embedding layer, an adaptive sampling layer, and a dynamic multi-resolution feature aggregation network that are connected in sequence.
[0042] Among them, the input layer is used to input the preprocessed log data.
[0043] The intelligent sampler is used to perform intelligent sampling on the input preprocessed log data. The action mechanism of the intelligent sampler is specifically described as follows:
[0044] First, a weight is calculated for each preprocessed log entry, and this weight reflects the importance of the log entry for fault prediction. The weight can be determined through learning of historical data. For example, by training a classification model (such as logistic regression, random forest, etc.) to estimate the probability of a fault occurring under different feature combinations. Let w i be the weight of the i-th log entry, then
[0045] w i = f(x i ; θ)
[0046] Among them, x i represents the feature vector of the i-th log entry, θ represents the model parameter, and f(·) represents the function used to calculate the weight.
[0047] Second, considering that the most recent log entries are often more relevant than older ones, a time decay factor α(t) can be introduced, and this factor decreases as the creation time t of the log entry increases. Specifically, the following exponential decay function can be used:
[0048] α(t) = e -λt
[0049] Among them, λ represents the hyperparameter that controls the decay rate, and t represents the time difference between the creation time of the log entry and the current time.
[0050] Third, based on the above two factors, the importance score of each log entry can be located as:
[0051] s i = w i × α(t i )
[0052] Among them, s i represents the comprehensive score that combines the feature weight and the time decay.
[0053] Fourth, in order to decide whether to select a certain log entry for further analysis, its importance score can be converted into a sampling probability p i , which is specifically expressed as follows:
[0054]
[0055] where p i represents the sampling probability of the i-th log entry.
[0056] Fifth, according to the calculated sampling probability p i , an actual sampling decision can be made using a random number generator. If the generated random number is less than p i , then select this log entry; otherwise skip it.
[0057] In summary, the intelligent sampler is represented as follows:
[0058]
[0059] where y i represents whether to select the i-th log entry (1 means select, 0 means not select), and the rand() function returns a uniformly distributed random number in the range [0, 1].
[0060] It should be noted that since the amount of log data is extremely large, directly using all log data as input to the model will cause excessive consumption of the model's computing resources. Through intelligent sampling, data points that best represent the set-top box failure behavior can be selectively retained, thereby reducing the data volume. In addition, there are often a large amount of redundant or duplicate information in log data, and this information is of no help for fault prediction and may even introduce noise. Sampling can help identify and remove this unnecessary information, enabling the model to focus more on meaningful data features.
[0061] In addition, it should be noted that the entire sampling process of the intelligent sampler is not static, but can dynamically change the sampling strategy according to the current analysis requirements, thereby reflecting the so-called "intelligent" nature. For example, when an abnormal behavior is predicted, the intelligent sampler will automatically increase the resolution during this time period to obtain more details; while in the normal state, the resolution will be reduced to save computing resources.
[0062] The hybrid embedding layer includes an input layer, a feature processing module, a concatenation layer, a fully connected layer, and an output layer connected in sequence. Among them, the input layer is used to receive the log data samples obtained by sampling with the intelligent sampler. The feature processing module is used to process the log data samples. Specifically, the feature processing module includes a normalization layer, an embedding matrix, and a pre-trained language model arranged in parallel. Among them, for the numerical features (such as CPU usage rate) in the log data samples, they are directly normalized through the normalization layer (the normalization layer can ensure that all features are within a similar distribution range, which is beneficial to subsequent feature fusion and cross-modal interaction operations); for the categorical features (such as error codes) in the log data samples, each category is mapped to a low-dimensional dense vector through the embedding matrix (by mapping the discrete categorical features into the low-dimensional dense vector space, these features can participate in calculations together with numerical features and retain the semantic relationships between categories); for the text features (such as the description information in the crash report) in the log data samples, the pre-trained language model (such as BERT) is used to generate the vector representation of the text (the pre-trained language model has been fully trained on a large-scale corpus, has good generalization ability, can effectively handle newly emerging log text data, and reduce the risk of overfitting). The concatenation layer is used to fuse the log data samples processed by the normalization layer, the embedding matrix, and the pre-trained language model to form a comprehensive feature vector. The fully connected layer is used to further compress the comprehensive feature vector into a fixed-length vector representation and output it through the output layer.
[0063] Exemplarily, for numerical features, the hybrid embedding layer first processes them through the normalization layer to ensure that all features are within a similar distribution range. For example, if the range of CPU usage rate is from 0 to 100%, and the memory usage rate is an absolute value in MB, normalization can adjust them to the same scale, such as within the range of [0, 1].
[0064] For categorical features, the embedding matrix is used to map each category to a low-dimensional dense vector. Suppose there is an error code field that contains multiple different error codes, and each error code can be represented by a vector of a predefined size. For example, the error code "E102" is mapped into a 5-dimensional vector [0.1, 0.3, -0.2, 0.6, 0.4].
[0065] For text features, the pre-trained language model (such as BERT) is used to generate the vector representation of the text. For example, for the description text in the application crash report: "The application unexpectedly closed while playing a video", a fixed-length vector representation can be obtained through the BERT model, and this vector captures the semantic information of this text.
[0066] Once all types of log data samples have been converted into vector form, the next step is to concatenate these vectors together to form a comprehensive feature vector. For example, suppose there are the following three types of features:
[0067] Numerical feature: Normalized CPU usage [0.8]
[0068] Categorical feature: Vector corresponding to the error code [0.1, 0.3, -0.2, 0.6, 0.4]
[0069] Textual feature: Vector from the BERT model [0.2, 0.5, -0.3,..., 0.7] (assumed to be a 768-dimensional vector)
[0070] These features will be directly concatenated to form a new vector, for example: [0.8, 0.1, 0.3, -0.2, 0.6, 0.4, 0.2, 0.5, -0.3,..., 0.7].
[0071] Through the above operations, it is possible to transform the original data into a form suitable for subsequent deep learning model processing while retaining all the information of the original data.
[0072] The hybrid embedding layer can form a richer feature representation by integrating different types of log data, which helps the model learn deeper patterns and thus improve the fault prediction accuracy of the set-top box.
[0073] The adaptive sampling layer includes a feature extractor, a scoring function, a sampling rate controller, and a sampling decision module connected in sequence. Among them, the feature extractor is used to extract relevant features from the processed log data samples output by the hybrid embedding layer and generate a new set of feature representations. The scoring function is used to calculate a score s for each log entry to reflect the value of the log entry for fault prediction. The scoring function is specifically expressed as follows:
[0074] s = g(f(x); β)
[0075] Where β represents the parameters of the scoring function, and g(·) represents the scoring model, such as logistic regression or random forest.
[0076] The sampling rate controller is used to dynamically adjust the sampling rate r to achieve adaptive sampling. The sampling rate controller is specifically expressed as follows:
[0077] r = h(s, S; φ)
[0078] Among them, S represents system state information, such as the current memory occupancy rate, CPU load, etc., φ represents the parameters of the sampling rate controller, and h(·) can be designed according to actual requirements. For example, it can be a rule-based method (if the score is higher than the threshold and the system resources are sufficient, the sampling rate is provided).
[0079] The sampling decision module determines whether to select a certain log entry according to the calculated sampling rate r. The sampling decision module is expressed as follows:
[0080]
[0081] Among them, rand() returns a uniformly distributed random number in the interval [0,1], and y represents a binary number indicating whether to select the i-th log entry (1 means select, 0 means not select).
[0082] The dynamic multi-resolution feature aggregation network includes a multi-scale input layer, a pyramid pooling module, an adaptive resolution adjuster, an integrated attention mechanism, a recurrent neural network (RNN), and an interpretability output interface connected in sequence. Among them, the multi-scale input layer can simultaneously process data from different time windows of the log, such as data streams including short time windows (such as 5 minutes), medium time windows (such as 1 hour), and long time windows (such as 24 hours). Each time window corresponds to a different resolution, so as to be able to capture short-term fluctuations, medium-term trends, and long-term patterns. The pyramid pooling module is used to decompose the input log data into representations of multiple scales. The pyramid pooling module contains several sub-modules, and each sub-module is responsible for performing max-pooling or average-pooling operations on a receptive field of a specific size. This can ensure that the model not only pays attention to local details but also understands global information.
[0083] The adaptive resolution adjuster can dynamically change the input resolution of some parts according to the current analysis requirements. For example, when abnormal behavior is predicted, the adaptive resolution adjuster will automatically increase the resolution during this period to obtain more details; while in the normal state, the resolution is reduced to save computing resources. The adaptive resolution adjuster is expressed as follows:
[0084]
[0085] Among them, X represents the input log data sequence, C represents the current scenario information, including but not limited to the abnormal behavior prediction result, system load status, r(C) represents the resolution adjustment factor determined by the scenario information C, Upsample(X,r) represents the operation of increasing the time resolution of the sequence X by r times, which can be achieved by linear interpolation for example, and Downsample(X,r) represents the operation of reducing the time resolution of the sequence X by r times, which can be achieved by max-pooling for example.
[0086] Among them, the resolution adjustment factor r(C) is expressed as follows:
[0087] r(C) = g(C; θ)
[0088] Among them, g represents a parameterized function used to predict the most appropriate resolution adjustment ratio according to the context information C, θ represents the parameter of the function g, r(C) > 1 indicates increasing the resolution, and r(C) < 1 indicates decreasing the resolution.
[0089] The adaptive resolution adjuster can dynamically change the input resolution of some parts according to the current analysis requirements. For example, when an abnormal behavior is predicted, the resolution is automatically increased during that time period to obtain more details; while in the normal state, the resolution is decreased to save computing resources.
[0090] The integrated attention mechanism enables the model to focus on those time periods or features that are most valuable for fault prediction. The integrated attention mechanism improves the importance of key information by learning to assign weights to feature points at different time and spatial positions.
[0091] The recurrent neural network (RNN) is used to process sequential log data to capture trends and dependencies evolving over time;
[0092] The interpretable output interface can generate an easy-to-understand report indicating which specific features or time periods are most likely to have caused the predicted fault risk.
[0093] Through the collaborative work of the above-mentioned modules, the dynamic multi-resolution feature aggregation network can efficiently process and deeply analyze the set-top box log data, not only accurately predicting potential faults of the set-top box, but also providing detailed explanations, thereby greatly improving the prediction performance of the model for set-top box faults.
[0094] The feature fusion network includes a multi-modal input layer, a modality-specific feature extraction layer, a cross-modal interaction layer, and a global context awareness module connected in sequence. Among them, the multi-modal input layer includes a log parsing module, a time series encoder, and a text embedding generator. The log parsing module is used to receive and parse the features of different types of log data extracted by the backbone network. The time series encoder applies LSTM or GRU to features with a time order to capture time dependencies and trends, and at the same time uses the self-attention mechanism in the Transformer architecture to enhance the model's understanding of long-term dependencies. The text embedding generator uses a pre-trained language model (such as BERT) to generate high-quality vector representations for features containing text information, or adopts a method based on the Bag of Words (BoW) model for simple and fast text feature extraction.
[0095] The modality-specific feature extraction layer includes a dedicated feature extractor, which, for example, includes a time series anomaly prediction enhancer, a text semantic understanding module, an abnormal behavior predictor, and a multi-dimensional health assessor.
[0096] Among them, the time series anomaly prediction enhancer includes an input layer, an encoder, a variational autoencoder, and a decoder connected in sequence. The input layer is used to receive time series data, such as network connection state changes and performance statistics data. The encoder includes an LSTM layer and a fully connected layer. The LSTM layer is used to capture long-term dependencies in the time series (multiple LSTM layers can be stacked to increase the depth and expressive power of the model, for example, 3 to 5 LSTM layers can be stacked), and the fully connected layer is used to convert the output of the LSTM into a latent variable. The spatial variational autoencoder is used to calculate the mean and variance of the latent variable based on the output of the fully connected layer, and perform reparameterization trick sampling according to the calculated mean and variance to generate a latent variable. The decoder includes a reverse fully connected layer and an LSTM layer. The reverse fully connected layer is used to map the latent variable back to the original data space, and the LSTM layer is used to reconstruct the time series data to ensure that the model can learn the data representation of the normal mode.
[0097] The text semantic understanding module includes an input layer, a pre-trained language model layer, and a domain-specific fine-tuning layer connected in sequence. Among them, the input layer is used to receive log data rich in text information, such as application crash reports, system error descriptions, etc. The pre-trained language model layer can, for example, use the BERT pre-trained language model as the basic architecture to perform preliminary encoding on the input text to obtain context-aware word embedding representations. The domain-specific fine-tuning layer includes a domain knowledge graph fusion layer and a task-specific fully connected layer. The domain knowledge graph fusion layer adjusts the weights of BERT or adds additional attention mechanisms by embedding domain-specific knowledge graphs into the model, enabling the model to better understand and process industry-specific terms and technical details. The task-specific fully connected layer further refines features and outputs the final text representation according to the specific fault prediction task requirements.
[0098] The abnormal behavior predictor includes an environment simulation layer, a reinforcement learning agent layer, a feature extraction layer, and a decision layer connected in sequence. Among them, the environment simulation layer can construct a simplified user interaction environment to simulate behaviors such as the user's remote control operations and channel switching. The reinforcement learning agent layer uses algorithms such as DQN (Deep Q-Network) or A3C (Asynchronous Advantage Actor-Critic) to let the agent learn in this environment to identify behavior patterns that may cause the set-top box to be unstable. The feature extraction layer includes a convolutional neural network and a long short-term memory network, which are respectively used to extract local features from historical behavior data and capture the time series characteristics of user behaviors to help predict future user behavior trends. The decision layer combines the results of reinforcement learning and the output of the feature extraction layer to give a judgment on whether there is a potential fault risk for the set-top box.
[0099] The multi-dimensional health evaluator includes an input layer, a multi-head self-attention mechanism layer, a global pooling layer, and a fully connected layer connected in sequence. Among them, the input layer is used to receive data streams from different sensors, including temperature readings, disk SMART status, etc. The multi-head self-attention mechanism layer allows the model to focus on the interactions between different dimensions of the input to effectively capture cross-modal correlations. The global pooling layer is used to summarize information from all dimensions to form a comprehensive health indicator. The fully connected layer is used to further refine the comprehensive health indicator and output the final health comment result. In addition, a feed-forward neural network is set after each head in the multi-head self-attention mechanism layer for non-linear transformation.
[0100] Since different types of log data include different semantic information, the present application selects different feature extractors for each type of log data, which helps the model to better understand this semantic information, so as to discover the potential correlations and interactions between modalities to form a richer feature space, and further provide information support in more dimensions for set-top box fault prediction, improving the accuracy and reliability of fault prediction.
[0101] The cross-modal interaction layer includes a multi-modal feature alignment module, a feature cross-connection layer, and a fusion strategy selector connected in sequence. Among them, the multi-modal feature alignment module includes a feature mapper and a batch normalization layer. The feature mapper is used to map features of different modalities into a common space so that features of different modalities can directly interact with each other. For example, a linear transformation can be used to adjust the dimension of each modal feature vector to match the target space. The batch normalization layer is used to perform a normalization operation on features of different modalities, making features of different modalities comparable. The feature cross-connection layer includes an adaptive weighting module, a cross-modal interaction module, and a batch normalization layer connected in sequence. The adaptive weighting module is used to calculate importance weights for the aligned multi-modal features based on the attention mechanism, enabling the model to focus on the features most critical to the current task. Specifically, the attention score of each modal feature can be calculated by the following formula:
[0102] s i = softmax(W a ·f i + b a )
[0103] where s i represents the attention score of the i-th modal feature, f i represents the feature vector corresponding to the attention score, W a represents a learnable circle matrix, and b a represents a learnable bias term.
[0104] The attention scores of each modal feature are weighted and summed for the original features to obtain a weighted feature representation f':
[0105]
[0106] The cross-modal interaction module is used to promote in-depth interaction between features of different modalities. It can adopt the way of residual connection or skip connection, allowing information exchange between high-level features and low-level features, and enhancing the expressiveness of the model. The cross-modal interaction module takes the weighted feature representation f' as input and combines it with the original feature f to form a new feature representation f'', which is specifically expressed as follows:
[0107] f” = f' + f
[0108] The batch normalization layer is used to normalize the features of different modalities after deep interaction, making the features of different modalities comparable and accelerating the training process of the model. The normalized multi-modal features are fed into the fusion strategy selector.
[0109] The feature cross-connection layer, by integrating the attention mechanism and residual or skip connections, can not only enhance the information flow between different modality features, but also improve the flexibility and robustness of the model in dealing with complex and heterogeneous data.
[0110] The fusion strategy selector is used to automatically select the most suitable fusion strategy according to different situations. The fusion strategies include, for example, simple addition, concatenation or more complex non-linear combination methods. The fusion strategy selector is expressed as follows:
[0111]
[0112] where, f i represents the i-th feature source, w i represents the weight of each feature source, z i represents the gating value of each feature source obtained through the soft gating mechanism, taking values in [0, 1], p j represents the probability distribution for selecting the fusion strategy, s j represents the j-th fusion strategy.
[0113] The fusion strategy selector allows the model to learn the optimal feature fusion method during the training process and can dynamically adjust the fusion strategy according to different input conditions, thereby improving the accuracy and robustness of the set-top box fault prediction.
[0114] The global context awareness module includes a context awareness pooling layer and a conditional random field. The context awareness pooling layer is used to combine the context information obtained from the outside (such as environmental temperature, power supply stability, etc.) and integrate local features into global features through weighted pooling operations, so that the final output can not only reflect the characteristics of a single log entry, but also reflect the impact of the overall environment on the operating state of the set-top box. The conditional random field (CRF) is applied to time series data to help the model understand the transition probability between adjacent time points, so that it can better capture the evolution process of continuous fault patterns, enabling the model to not only identify immediate abnormal situations, but also predict potential long-term problems, thereby greatly improving the accuracy and reliability of the model for set-top box fault prediction.
[0115] The output network is used to receive the comprehensive features processed by the feature fusion network and convert them into specific fault prediction results. The output network includes a fully connected layer, a Dropout layer, a batch normalization layer, and an activation function connected in sequence. Among them, the fully connected layer is used to receive the high-dimensional feature vector from the feature fusion network and compress the high-dimensional features into a lower-dimensional space through a series of linear transformations and non-linear activation functions (such as ReLU, Leaky ReLU, etc.) to reduce redundant information and extract more representative feature representations. The Dropout layer is used to randomly discard a certain proportion of neurons during training to prevent overfitting and improve generalization ability. The normalization layer is used to normalize each batch of data to accelerate the training process and help stabilize the learning process. The activation function uses the Softmax activation function to output the probability values of each fault category of the set-top box.
[0116] In another exemplary embodiment, in step S300, the set-top box fault prediction model is trained through the following steps:
[0117] S310: Obtain the historical fault data of the set-top box and preprocess it, and divide the preprocessed historical fault data into a training set and a validation set according to a ratio, for example, the division ratio is 7:3;
[0118] S320: Set the training parameters. For example, the booster default selects the tree-based model gbtree, the learning_rate is set to 0.01, the max_depth is set to 3, and the maximum number of iterations is set to 300. Use the training set to train the model. During the model training process, when the number of model iterations meets the preset value, the model training is completed;
[0119] S330: Use the validation set to verify the trained model. During the verification process, if the mean absolute error (MAE), mean square error (MSE), and root mean square error (RMSE), which are used as model performance evaluation indicators, are all less than the thresholds (the threshold of MAE is set to 0.05, and the threshold of MSE is set to 0.001), then the model verification passes; otherwise, adjust the training parameters (for example, the learning_rate can be adjusted to 0.005, and the max_depth can be adjusted to 5) or expand the training set samples (for example, adjust the division ratio to 8:2) and retrain the model until the verification passes.
[0120] Next, a set of log data of a certain set-top box during a certain period is collected in this application to exemplarily illustrate the method described in this application. The log data specifically includes:
[0121] 1. System error and warning information:
[0122] Timestamp: January 15, 2025, 14:03:22; Error Code: E102 - Out of Memory Warning.
[0123] Timestamp: January 18, 2025, 09:15:45; Error Code: W204 - Network Connection Unstable Prompt.
[0124] 2. Application Crash Reports:
[0125] Application Name and Version: MediaCenter v3.2.1; Crash Time: January 16, 2025, 21:30:12; Crash Frequency: A total of 3 crashes occurred in the past week.
[0126] Application Name and Version: VideoPlayer v1.5.0; Crash Time: January 20, 2025, 18:45:07; Crash Frequency: Two crashes occurred in the past three days.
[0127] 3. Network Connection Status Changes:
[0128] Connection / Disconnection Time: January 17, 2025, 08:00:00 (Connected); January 17, 2025, 08:05:30 (Disconnected).
[0129] Reason for Connection Failure: DNS resolution failed when attempting to connect on January 22, 2025.
[0130] 4. Hardware Health Metrics:
[0131] Temperature Reading: The CPU temperature reached a peak of 85°C on January 21, 2025, and exceeded 75°C multiple times in the past week.
[0132] Disk SMART Status: The hard drive SMART attributes show uncorrected read errors, which may indicate poor hard drive health.
[0133] 5. User Behavior Patterns:
[0134] Remote Control Operation Frequency: Frequent volume adjustments and channel changes occurred on January 19, 2025, with the number of operations increasing by approximately 50% compared to normal days.
[0135] Playback Interruption: During video playback, there were three loading delays on January 20, 2025, each lasting approximately 15 seconds; and frequent buffering occurred.
[0136] 6. Performance Statistics:
[0137] CPU Usage Peak: Around 22:15 on January 21, 2025, the CPU usage soared to 98% and remained at that level for approximately 10 minutes.
[0138] Memory Occupancy Peak: During the same time period, the memory usage also reached a peak value close to 90%.
[0139] By analyzing the above data through the set-top box fault prediction model provided in this application, the following conclusions can be drawn:
[0140] 1. Insufficient Memory Problem: The E102 error code indicates that there is an insufficient memory problem with the set-top box. Combining the CPU and memory occupancy peaks and the crash frequency of the MediaCenter application, this may mean that the system encounters resource bottlenecks when processing multitasks or running large applications. It is recommended to check if there are any unnecessary background processes consuming resources and consider whether additional physical memory is needed.
[0141] 2. Unstable Network Connection: The W204 warning prompt and DNS resolution failure indicate that the set-top box may have encountered network configuration problems or the unavailability of external network services. Considering the frequent disconnection of the network connection and the crashes of the VideoPlayer application (which may be related to network playback), the local network settings, router configuration, and the status of the Internet service provider should be checked.
[0142] 3. Hardware Health Problem: The overheating of the CPU may be caused by the low efficiency of the cooling system or the blockage of the heat dissipation channels, or it may also be due to long-term high-load operation. High temperature may lead to a decline in system performance or even hardware damage, so immediate measures should be taken to cool down. The disk SMART attribute showing uncorrected read errors is a serious signal indicating that the hard disk may have physical damage or other serious problems. Given the importance of the hard disk for data storage, it is recommended to back up important data as soon as possible and consider replacing the hard disk.
[0143] 4. User Behavior Patterns: The increased frequency of remote control operations and playback interruptions reflect that users have experienced a decline in the quality of service (QoS). Although this is not a direct hardware or software failure, it does point to potential service performance problems, which may be caused by the above-mentioned resource shortages or network problems.
[0144] 5. Performance Statistics: The peak values of CPU usage and memory occupancy further confirm the tight situation of system resources. This not only affects the response speed of the system but may also indirectly lead to the instability of application programs (such as crashes).
[0145] In summary, the set-top box currently faces multiple problems, including but not limited to insufficient memory, unstable network connection, overheating of hardware, and poor disk health.
[0146] To address the above problems existing in the set-top box, maintenance personnel can consider taking the following measures:
[0147] Optimize the system and applications to reduce resource consumption;
[0148] Check and fix network configuration issues;
[0149] For hardware problems, especially disk health, timely evaluation and necessary replacement should be carried out;
[0150] Monitor and control the usage of CPU and memory to ensure that they do not remain in a high-load state for a long time;
[0151] Upgrade the hardware (such as memory modules) to improve the overall performance of the set-top box.
[0152] In another exemplary embodiment, the present application further provides a set-top box fault prediction device based on log data. The device includes: an acquisition module 100 for acquiring set-top box log data within a set period, where the set-top box log data includes system error and warning information, application crash reports, network connection status changes, hardware health indicators, user behavior patterns, and performance statistics; a preprocessing module 200 for preprocessing the log data; a model construction and training module 300 for constructing a set-top box fault prediction model and training the model; and a prediction model 400 for inputting the preprocessed log data into the trained set-top box fault prediction model to predict the faults of the set-top box.
[0153] Optionally, the preprocessing module 200 includes: a cleaning sub-module for cleaning the log data; a formatting sub-module for formatting the cleaned log data; and a feature engineering sub-module for performing feature engineering on the formatted log data.
[0154] Based on the above embodiments, with reference to Figure 4 , a computer-readable storage medium of an exemplary embodiment of the present application is described. Please refer to Figure 4 , the shown computer-readable storage medium is an optical disc 40, on which a computer program (i.e., a program product) is stored. When the computer program is run by a processor, it will implement the steps recorded in the above method embodiments, for example, acquiring set-top box log data within a set period; preprocessing the log data; constructing a set-top box fault prediction model and training the model; inputting the preprocessed log data into the trained set-top box fault prediction model to predict the faults of the set-top box. The specific implementation methods of each step will not be repeated here.
[0155] It should be noted that the computer-readable storage medium includes, but is not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other optical and magnetic storage media, which will not be elaborated here one by one.
[0156] Based on the above embodiments, the present application further provides an electronic device. The following refers to Figure 5 to describe the electronic device for file downloading in the exemplary embodiments of the present application.
[0157] Figure 5 The block diagram of an exemplary electronic device 50 suitable for implementing the embodiments of the present application is shown. The electronic device 50 may be a computer system or a cloud server. Figure 5 The shown electronic device 50 is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.
[0158] As Figure 5 shown, the electronic device 50 includes, but is not limited to: one or more processors or processing units 501, a system memory 502, and a bus 503 connecting different system components (including the system memory 502 and the processing unit 501).
[0159] The electronic device 50 typically includes various computer system-readable media. These media can be any available media accessible by the electronic device 50, including volatile and non-volatile media, removable and non-removable media.
[0160] The system memory 502 may include computer system-readable media in the form of volatile memory, such as random access memory (RAM) 5021 and / or cache memory 5022. The electronic device 50 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, ROM 5023 may be used to read and write non-removable, non-volatile magnetic media ( Figure 5 not shown, commonly referred to as a "hard disk drive"). Although not shown in Figure 5As shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., "floppy disk") and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., CD-ROM, DVD-ROM or other optical media) can be provided. In these cases, each drive can be connected to the bus 503 through one or more data medium interfaces. The system memory 502 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of the present application.
[0161] A program / utility 5025 having a set (at least one) of program modules 5024 can be stored, for example, in the system memory 502, and such program modules 5024 include but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment. The program modules 5024 generally perform the functions and / or methods in the embodiments described in the present application.
[0162] The electronic device 50 can also communicate with one or more external devices 504 (such as a keyboard, a pointing device, a display, etc.). Such communication can be carried out through the input / output (I / O) interface 505. And, the electronic device 50 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN) and / or a public network, such as the Internet) through the network adapter 506. As Figure 5 shown, the network adapter 506 communicates with other modules (such as the processing unit 501, etc.) of the electronic device 50 through the bus 503. It should be understood that although Figure 5 not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 50.
[0163] The processing unit 501 executes various functional applications and data processing by running the programs stored in the system memory 502. For example, it obtains the set-top box log data within a set period; preprocesses the log data; constructs a set-top box fault prediction model and trains the model; inputs the preprocessed log data into the trained set-top box fault prediction model to predict the faults of the set-top box. The specific implementation manners of each step are not repeated here. It should be noted that although several units / modules or sub-units / sub-modules of the file concurrent download device are mentioned in the above detailed description, this division is only exemplary and not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more units / modules described above can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided and embodied by multiple units / modules.
[0164] In the description of the present application, it should be noted that the terms "first", "second", and "third" are only used for descriptive purposes and should not be construed as indicating or implying relative importance.
[0165] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0166] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical functional division, and there can be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some communication interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0167] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0168] In addition, the functional units in each embodiment of the present application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0169] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a cloud server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs and other various media that can store program codes.
[0170] The above are only the preferred embodiments of the present application, which do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present application.
Claims
1. A set-top box fault prediction method based on log data, characterized in that: The method comprises: Obtaining set-top box log data within a set period, wherein the set-top box log data includes system error and warning information, application crash reports, network connection status changes, hardware health indicators, user behavior patterns, and performance statistics; Preprocessing the log data; Constructing a set-top box fault prediction model and training the model; the set-top box fault prediction model introduces an intelligent sampler to intelligently sample the pre-processed log data to reduce the amount of data; The pre-processed log data is input into a trained set-top box fault prediction model to predict the fault of the set-top box.
2. A set-top box fault prediction method based on log data according to claim 1, characterized in that: Preprocessing the log data includes: Cleaning the log data; Formatting the cleaned log data; Feature engineering is performed on the formatted log data.
3. A set-top box fault prediction method based on log data according to claim 1, characterized in that: The set-top box failure prediction model includes: Backbone network, feature fusion network and output network, where: The backbone network is used to extract features from the preprocessed log data; The feature fusion network is used to fuse the features extracted by the backbone network to obtain comprehensive features; The output network is used to receive the comprehensive features and output the fault prediction result of the set-top box.
4. A set-top box fault prediction method based on log data according to claim 3, characterized in that: The backbone network includes: Input layer, smart sampler, hybrid embedding layer, adaptive sampling layer, and dynamic multi-resolution feature aggregation network.
5. The method for predicting set-top box failure based on log data according to claim 3, characterized in that: The feature fusion network includes the following connected in sequence: Multimodal input layer, modality-specific feature extraction layer, cross-modal interaction layer, and global context-aware module.
6. A set-top box fault prediction method based on log data according to claim 3, characterized in that: The output network includes: Fully connected layers, Dropout layers, Batch Normalization layers, and activation functions.
7. A set-top box fault prediction method based on log data according to claim 1, characterized in that: The set-top box fault prediction model is trained by the following steps: Obtain and preprocess historical fault data of the set-top box, and divide the preprocessed historical fault data into a training set and a validation set in proportion; Set training parameters and use the training set to train the model. During the model training process, when the number of model iterations meets the preset value, the model training is completed; The trained model is verified using the validation set. During the verification process, if the mean absolute error, mean square error, and root mean square error, which are the model performance evaluation indicators, are all less than the threshold, the model verification is passed; otherwise, the training parameters are adjusted or the training set samples are expanded to retrain the model until the verification is passed.
8. A set-top box fault prediction device based on log data, characterized in that: The device comprises: An acquisition module, used to acquire set-top box log data within a set period, wherein the set-top box log data includes system error and warning information, application crash reports, network connection status changes, hardware health indicators, user behavior patterns, and performance statistics; A preprocessing module, used for preprocessing the log data; Model building and training module, used to build set-top box failure prediction model and train the model; The prediction model is used to input the pre-processed log data into a trained set-top box fault prediction model to predict the set-top box fault.
9. A storage medium, characterized in that: The method comprises instructions, which, when executed on a computer, enable the computer to execute the method according to any one of claims 1 to 7.
10. An electronic device, characterized in that: The electronic device comprises: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Video abnormal link micro-segmentation positioning method and system
CN122205073A
Video abnormal link micro-segmentation positioning method and system
CN122205073B