Lag detection method and device, electronic equipment and storage medium
By analyzing n-frame images using deep learning network models on the mobile terminal, the problem that mobile terminal lag detection relies on user subjective perception is solved, and more accurate and extensive lag detection is achieved, reducing labor costs.
Patent Information
- Application Number
- CN202311511596.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-13
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, mobile terminal lag detection is relatively lacking, mainly relying on the user's subjective perception, there are problems of misjudgment and inconvenient positioning, and common methods cannot detect unresponsive pop-ups, and the detection results are inaccurate.
By acquiring n-frame images, calling the deep learning network model for lag detection, using the first model and the second model for similarity detection and feature extraction, respectively, to determine whether there is lag.
It realizes more accurate lag detection, improves the accuracy and range of detection, reduces labor costs, can automatically capture logs and record videos, and saves human resources.
Smart Images

Figure CN119992397A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of deep learning, and in particular to a freeze detection method, device, electronic device and storage medium. Background Art
[0002] Among the related technologies, there are relatively few technologies for detecting freezes on mobile terminals. The judgment of problems such as page unresponsiveness, frame drops, and frame loss is more dependent on the user's subjective perception, which may lead to certain misjudgments and inconvenient positioning. Common freeze detection schemes mainly include the use of time period division method, pixel timestamp conversion method, two-frame difference image grayscale method, etc. for freeze detection, but there are problems such as failure to detect unresponsive pop-up windows and inaccurate detection results. Summary of the invention
[0003] In order to overcome the problems existing in the related art, the present disclosure provides a freeze detection method, device, electronic device and storage medium.
[0004] According to a first aspect of an embodiment of the present disclosure, a jam detection method is provided, including:
[0005] Obtain n frames of images, wherein the n frames of images include a current frame of a terminal display screen, and n is a positive integer greater than or equal to 2; call a network model for jamming detection, and obtain a jamming detection result based on the network model and the n frames of images; wherein the jamming detection result includes the presence of jamming or the absence of jamming.
[0006] In one embodiment, the network model includes a first model and a second model, the first model is used to obtain a jamming detection result by performing a similarity test on the n frames of images, and the second model is used to obtain a jamming detection result by performing a feature extraction on the current frame and performing a similarity test between the processed features and the preset features; obtaining the jamming detection result based on the network model and the n frames of images includes: determining that jamming occurs in the n frames of images in response to the first model or the second model, and determining that the jamming detection result is that jamming exists.
[0007] In one embodiment, the network model includes a first model; the jamming detection result is obtained based on the network model and the n frames of images, including: obtaining feature vectors corresponding to n frames of images, and arranging the feature vectors corresponding to the n frames of images in chronological order to obtain a first sequence; subtracting the feature vectors of two adjacent frames in the first sequence to obtain a second sequence containing n-1 difference vectors; inputting the second sequence into the first model to obtain the predicted value of the first model for the difference vector between the current frame and the frame before the current frame; determining the predicted value of the current frame based on the predicted value; comparing the predicted value of the current frame with the current frame for similarity; and determining that jamming exists in response to the similarity being less than a first threshold.
[0008] In one implementation, the network model includes a second model; and obtaining a freeze detection result based on the network model and the n frames of images includes:
[0009] Acquire a feature vector of the current frame; perform at least one convolution operation and pooling operation on the feature vector using a convolution kernel of a preset size to obtain a one-dimensional feature vector; and compare the one-dimensional feature vector with the preset feature for similarity;
[0010] In response to the similarity being greater than a second threshold, it is determined that a freeze exists.
[0011] In one implementation, the first model is trained in the following manner:
[0012] A training video and a test video are obtained, and the training video and the test video are respectively segmented to obtain i training frames and j test frames, where i and j are positive integers; feature extraction is performed on the i training frames and the j test frames to obtain feature vectors corresponding to the i training frames and the j test frames respectively; the feature vectors of the i training frames are subtracted to obtain difference vectors of i-1 training frames; the feature vectors of the j test frames are subtracted to obtain difference vectors of j-1 test frames; the difference vectors of the i-1 training frames are input into a first prediction model for training, and the first prediction model is constrained based on the difference vectors of the j-1 test frames to obtain a first model through training.
[0013] In one implementation, the second model is trained in the following manner:
[0014] The feature vectors corresponding to the sample images are marked, the feature vectors corresponding to the images without pop-up windows or with non-stuttering pop-up windows are marked as positive samples, and the feature vectors of the images with stuttering pop-up windows are input into the second model as negative samples; the feature vectors are subjected to feature extraction through multiple convolutional layers and pooling layers to obtain the extracted feature vectors; based on the extracted feature vectors, the model parameters in the second model are adjusted; the above process is repeated until the recognition result of the second model meets the preset effect to obtain the second model.
[0015] In one embodiment, the method further comprises:
[0016] In response to the jam detection result indicating that jamming exists, at least one of the following operations is performed: capturing a log containing jamming information in the terminal; saving image information of the current frame; and saving video information containing the current frame.
[0017] According to a second aspect of an embodiment of the present disclosure, a jam detection device is provided, including:
[0018] An acquisition unit is used to acquire n frames of images, wherein the n frames of images include a current frame of a terminal display screen, and n is a positive integer greater than or equal to 2; a processing unit is used to call a network model for performing a jamming detection, and obtain a jamming detection result based on the network model and the n frames of images; wherein the jamming detection result includes the presence of a jamming or the absence of a jamming.
[0019] In one embodiment, the network model for detecting stuttering includes a first model and a second model, the first model being used to obtain a stuttering detection result by performing a similarity test on the n frames of images, and the second model being used to obtain a stuttering detection result by performing a feature extraction on the current frame and performing a similarity test between the processed features and preset features; the processing unit obtains the stuttering detection result based on the network model and the n frames of images in the following manner: in response to determining that stuttering occurs in the n frames of images by the first model or the second model, the stuttering detection result is determined to be the presence of stuttering.
[0020] In one embodiment, the network model includes a first model; the processing unit obtains a freeze detection result based on the network model and the n frames of images in the following manner: obtain feature vectors corresponding to n frames of images, and arrange the feature vectors corresponding to the n frames of images in chronological order to obtain a first sequence; subtract feature vectors of two adjacent frames in the first sequence to obtain a second sequence containing n-1 difference vectors; input the second sequence into the first model to obtain a predicted value of the first model for the difference vector between the current frame and the frame before the current frame; determine a predicted value of the current frame based on the predicted value; compare the predicted value of the current frame with the current frame for similarity; and determine that freeze exists in response to the similarity being less than a first threshold.
[0021] In one embodiment, the network model includes a second model; the processing unit obtains a jamming detection result based on the network model and the n frames of images in the following manner: obtaining a feature vector of the current frame; performing at least one convolution operation and pooling operation on the feature vector using a convolution kernel of a preset size to obtain a one-dimensional feature vector; comparing the one-dimensional feature vector with the preset feature for similarity; and determining that jamming exists in response to the similarity being greater than a second threshold.
[0022] In one implementation, the first model is trained in the following manner:
[0023] A training video and a test video are obtained, and the training video and the test video are respectively segmented to obtain i training frames and j test frames, where i and j are positive integers; feature extraction is performed on the i training frames and the j test frames to obtain feature vectors corresponding to the i training frames and the j test frames respectively; the feature vectors of the i training frames are subtracted to obtain difference vectors of i-1 training frames; the feature vectors of the j test frames are subtracted to obtain difference vectors of j-1 test frames; the difference vectors of the i-1 training frames are input into a first prediction model for training, and the first prediction model is constrained based on the difference vectors of the j-1 test frames to obtain a first model through training.
[0024] In one implementation, the second model is trained in the following manner:
[0025] The feature vectors corresponding to the sample images are marked, the feature vectors corresponding to the images without pop-up windows or with non-stuttering pop-up windows are marked as positive samples, and the feature vectors of the images with stuttering pop-up windows are input into the second model as negative samples; the feature vectors are subjected to feature extraction through multiple convolutional layers and pooling layers to obtain the extracted feature vectors; based on the extracted feature vectors, the model parameters in the second model are adjusted; the above process is repeated until the recognition result of the second model meets the preset effect to obtain the second model.
[0026] In one embodiment, the method further comprises:
[0027] In response to the jam detection result indicating that jamming exists, at least one of the following operations is performed: capturing a log containing jamming information in the terminal; saving image information of the current frame; and saving video information containing the current frame.
[0028] According to a third aspect of an embodiment of the present disclosure, there is provided an electronic device, including:
[0029] A processor; a memory for storing processor executable instructions; wherein the processor is configured to: execute the jamming detection method described in the first aspect or any one of the embodiments of the first aspect.
[0030] According to a fourth aspect of an embodiment of the present disclosure, a storage medium is provided, in which instructions are stored. When the instructions in the storage medium are executed by a processor of a terminal, the terminal can perform the method described in the first aspect or any one of the implementations of the first aspect.
[0031] The technical solution provided by the embodiments of the present disclosure may include the following beneficial effects: by inputting n frames of images contained in the terminal display screen into the network model, the n frames of images containing the current frame are analyzed and processed based on the network model to obtain more accurate jamming detection results, improve the accuracy of jamming detection, reduce labor costs, and increase the jamming detection range.
[0032] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0034] Figure 1 The figure is a flow chart of a jamming detection method according to an exemplary embodiment.
[0035] Figure 2 The figure is a flow chart of a jamming detection method according to an exemplary embodiment.
[0036] Figure 3 The figure is a flowchart of a first model jamming detection method according to an exemplary embodiment.
[0037] Figure 4 is a flowchart of a second model jam detection method according to an exemplary embodiment
[0038] Figure 5 is a flowchart of a first model training method according to an exemplary embodiment.
[0039] Figure 6 is a flowchart of a second model training method according to an exemplary embodiment.
[0040] Figure 7 The figure is a schematic diagram of a jamming detection scenario according to an exemplary embodiment.
[0041] Figure 8 is a flowchart of training a first model according to an exemplary embodiment.
[0042] Fig. 9 is a flowchart of using a first model according to an exemplary embodiment.
[0043] Fig.10 is a schematic diagram showing the use of a second model structure according to an exemplary embodiment.
[0044] Fig.11 is a schematic diagram of a training process using a second model according to an exemplary embodiment.
[0045] Fig.12 is a schematic diagram showing the change in accuracy of a training process using a second model according to an exemplary embodiment.
[0046] Fig.13 It is a schematic diagram showing the change of loss value in the training process using the second model according to an exemplary embodiment.
[0047] Fig.14 It is a schematic diagram showing the change of training effect using the first model according to an exemplary embodiment.
[0048] Fig.15 is a schematic diagram of a normal pop-up window according to an exemplary embodiment.
[0049] Fig.16 The figure is a schematic diagram of a stuck pop-up window according to an exemplary embodiment.
[0050] Fig.17 The figure is a block diagram of a jamming detection device according to an exemplary embodiment.
[0051] Fig.18 The invention is a block diagram of a device for jamming detection according to an exemplary embodiment. DETAILED DESCRIPTION
[0052] Here, exemplary embodiments will be described in detail, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present disclosure.
[0053] Deep learning has been widely used in the field of computer vision, and convolutional neural networks are also one of the more important research directions in this field. Convolutional neural networks are widely used in image classification, target detection and other fields.
[0054] Currently, there are few technologies related to mobile terminal freeze detection. The judgment of problems such as page unresponsiveness and frame drops relies heavily on the user's subjective perception, which may lead to certain misjudgments and inconvenient positioning.
[0055] In view of this, the embodiments of the present disclosure focus on the field of Android mobile terminals, and use deep learning technology in the detection of freezes on mobile phones, aiming to solve the problem of freezes and unresponsiveness on mobile terminals. The embodiments of the present disclosure implement freeze detection on the operation of terminal devices based on multi-frame images, and detect freezes and unresponsiveness on the mobile phone, and record and save corresponding error reports and recordings of problem videos. In addition, the freeze detection method in the embodiments of the present disclosure can effectively reduce costs, free up manual work for stress testing, help R&D to locate freeze anomalies, and handle problems with high quality, which can bring greater benefits.
[0056] In the disclosed embodiment, the electronic device automatically controls the terminal device through a wired connection or a wireless connection, and manipulates the terminal device to run a specific test application (Application, APP). In the specific test application, it determines whether the current terminal has a lag phenomenon by running a trained network model. If it is determined that a lag phenomenon exists, a problem report is captured; if no lag phenomenon exists, detection continues.
[0057] Figure 1 is a flow chart of a jam detection method according to an exemplary embodiment. Figure 1 As shown, the following steps are included.
[0058] In step S11, n frames of images are acquired.
[0059] In the embodiment of the present disclosure, the n frames of images include the current frame of the terminal display screen, and n is a positive integer greater than or equal to 2. Since the n frames of images must include the current frame image, and n is a positive integer greater than or equal to 2, the embodiment of the present disclosure can obtain the image of the current frame and the image of at least one other frame.
[0060] In the disclosed embodiment, n frames of images including the current frame in the display screen can be acquired by an application running in the terminal within a preset time period, or can be acquired by uploading a recorded video to be detected and segmenting the video to be detected.
[0061] In step S12, a network model for jamming detection is called, and a jamming detection result is obtained based on the network model and n frames of images.
[0062] In the disclosed embodiment, the jamming detection result includes the presence of jamming or the absence of jamming.
[0063] In the disclosed embodiment, based on the acquired n frames of images, the n frames of images are input into a network model, and the network model performs analysis based on the images of all frames to determine whether there is any jamming in the n frames of images based on the analysis results, thereby improving the accuracy of jamming detection and expanding the scope of jamming detection.
[0064] In an exemplary embodiment, the 50 most recently displayed image frames on the terminal display screen can be obtained, one of which is the current frame. The 50 image frames are input into the network model, and the network model obtains a freeze detection result for the current frame based on operations such as feature extraction and prediction of the image content.
[0065] In the disclosed embodiment, the network model may be composed of one or more models, for example, a convolutional neural network model (CNN) or a long short-term memory network (LSTM). By using different types of models to process n frames of images, the judgment of freeze detection can be improved.
[0066] Figure 2 is a flow chart of a jam detection method according to an exemplary embodiment. Figure 2 As shown, the following steps are included.
[0067] In the disclosed embodiment, the network model includes a first model and a second model. The first model is used to obtain a jamming detection result by performing a similarity test on n frames of images, and the second model is used to extract features from the current frame and perform a similarity test between the processed features and the preset features to obtain a jamming detection result.
[0068] In step S21, n frames of images are acquired.
[0069] because Figure 2 The steps in step S21 in Figure 1 The process is the same as step S11 in , which will not be described in detail here, and reference may be made to the relevant description of the above embodiment.
[0070] In step S22, in response to the first model or the second model determining that a freeze occurs in the n frames of images, the freeze detection result is determined to be the presence of a freeze.
[0071] In the disclosed embodiment, the first model and the second model can be used to detect n frames of images at the same time, and the first model and the second model are ORed, that is, when a model detects that a jam occurs in n frames of images, the jam detection result is determined to be the presence of a jam. It should be understood that the first model and the second model belong to different types of network models. For example, the first model may be a convolutional neural network model, and the second model may be a long short-term memory network model. By performing jam detection using multiple models, it is possible to detect different types of jam situations and increase the jam detection range.
[0072] In the disclosed embodiment, the first model can be used to predict the current frame based on changes between adjacent frames, and whether there is a freeze condition can be determined based on the predicted value of the current frame and the current frame.
[0073] Figure 3 is a flowchart of a first model jam detection method according to an exemplary embodiment. Figure 3 As shown, the following steps are included.
[0074] In step S31, feature vectors corresponding to n frames of images are obtained, and the feature vectors corresponding to the n frames of images are arranged in chronological order to obtain a first sequence.
[0075] In the disclosed embodiments, the image may be preprocessed such as image scaling, grayscale conversion, noise reduction, etc. before extracting the feature vector to facilitate the feature extraction operation.
[0076] In the embodiments of the present disclosure, image processing technology can be used to extract features from images, convert image information into vector information, and determine the feature vector corresponding to each frame of the image. It should be understood that a variety of feature extraction methods can be used to obtain the feature values corresponding to the image, and the present disclosure does not limit the method for extracting feature vectors.
[0077] In the disclosed embodiment, the feature vectors corresponding to n frames of images are arranged in time order, and the obtained time sequence is set as a first sequence, wherein adjacent feature vectors in the first sequence correspond to images of adjacent frames in time.
[0078] In step S32, the feature vectors of two adjacent frames in the first sequence are subtracted to obtain a second sequence containing n-1 difference vectors.
[0079] In the embodiments of the present disclosure, the difference method is a method for converting a time series data set. In short, in a series of data, the difference between two adjacent values is obtained to obtain the transformation amount of two adjacent values. The difference can be performed by subtracting the previous observation from the current observation. The specific formula is as follows:
[0080] difference(t)=observation(t)-observation(t-1)
[0081] In the above formula, the parameter t represents time, observation(t) represents the feature vector of the frame corresponding to time t, difference(t) represents the difference vector obtained by subtracting the feature vectors corresponding to time t and time t-1 in the first sequence, and the obtained difference vectors are arranged in chronological order and set as the second sequence. It should be understood that since the difference vector needs to be obtained based on the difference between two adjacent feature vectors, only n-1 difference vectors can be obtained based on n feature vectors, that is, the difference vector corresponding to the feature vector at the initial moment cannot be obtained. It should be understood that the embodiment of the present disclosure can set the difference vector corresponding to the feature vector at the initial moment to a full zero vector, so that the feature vector at the initial moment has a corresponding difference vector.
[0082] In step S33, the second sequence is input into the first model to obtain the predicted value of the first model for the difference vector between the current frame and the frame before the current frame.
[0083] In the disclosed embodiment, the difference vector in the second sequence is input into the first model, and the first model processes the content in the second sequence to obtain the predicted value of the difference vector between the current frame and the frame before the current frame. For example, the difference vector corresponding to two adjacent frames from the second frame to the 99th frame is obtained through the feature vector corresponding to the image of the first frame to the 99th frame, and the difference vector between the 99th frame and the 100th frame is determined based on the difference vector.
[0084] In step S34, based on the predicted value, the predicted value of the current frame is determined.
[0085] In the disclosed embodiment, the inverse difference method can be used to determine the prediction value of the current frame based on the feature vector of the previous frame and the prediction value of the difference vector between the current frame and the previous frame. The formula of the inverse difference method is as follows:
[0086] inverted(t)=difference(t)+observation(t-1)
[0087] In the above formula, parameter t represents time, observation(t-1) represents the feature vector of the frame before the frame at time t, difference(t) represents the difference vector between the feature vector of the frame at time t and the feature vector of the frame at time t-1, and inverted(t) represents the predicted value of the feature vector of the frame at time t.
[0088] In step S35, the predicted value of the current frame is compared with the current frame for similarity.
[0089] In the disclosed embodiment, a new vector can be obtained by obtaining the difference between the predicted value of the current frame and the current frame, and the new vector can be input into the fully connected layer to obtain a scalar, which can be used to represent the similarity between the predicted value of the current frame and the current frame.
[0090] In step S36, in response to the similarity being less than the first threshold, it is determined that there is a freeze.
[0091] In the disclosed embodiment, a sigmoid function can be used to map a scalar representing the degree of similarity between the predicted value of the current frame and the current frame to 0 or 1. For example, assuming that the first threshold is 0.5, if the scalar is greater than or equal to 0.5, it means that the predicted value of the current frame is similar to the current frame, and a mapping value of 1 is obtained, indicating that there is no jamming; if the scalar is less than 0.5, it means that the predicted value of the current frame is not similar to the current frame, and a mapping value of 0 is obtained, indicating that there is jamming.
[0092] It should be understood that other methods may be used in the embodiments of the present disclosure to perform similarity comparison and determine whether there is a freeze.
[0093] In an exemplary embodiment, assuming that a frame of image is acquired every 1 millisecond (ms), 50 frames of image can be acquired within 50ms, wherein the 50th frame of image is the current frame of image displayed by the terminal. Feature extraction is performed by convolutional neural network, and feature vectors corresponding to 50 frames of image are obtained respectively, and these feature vectors are arranged in chronological order to obtain a feature vector sequence, i.e., the first sequence. By calculating the difference vector between the 1st frame and the 2nd frame in the first sequence, the difference vector between the 2nd frame and the 3rd frame, and so on, until the difference vector between the 48th frame and the 49th frame from the last is obtained, the sequence of 48 difference vectors obtained is used as the second sequence. The second sequence is input into the pre-trained LSTM to obtain the predicted value of the difference vector between the current frame and the previous frame, i.e., the predicted value of the difference vector between the 49th frame and the 50th frame, and the predicted value of the feature vector corresponding to the 50th frame of image can be obtained by calculating the predicted value of the difference vector between the 49th frame and the 50th frame and the sum vector of the feature vector of the 49th frame. The predicted value of the feature vector corresponding to the 50th frame is compared with the feature vector corresponding to the 50th frame image for similarity. If the similarity is greater than the preset threshold, the stuttering detection result is determined to be that there is no stuttering; if the similarity is less than the preset threshold, the stuttering detection result is determined to be that there is stuttering.
[0094] In the disclosed embodiment, a convolutional neural network can be used to extract features of the current frame content, and the extracted feature vector is compared with a preset feature vector for similarity to detect freezes or unresponsive pop-up phenomena.
[0095] Figure 4 is a flow chart of a second model jam detection method according to an exemplary embodiment. Figure 4 As shown, the following steps are included.
[0096] In step S41, a feature vector of the current frame is obtained.
[0097] In the disclosed embodiment, by converting the graphics of the current frame into a matrix of a specific size, the graphics of the previous frame are preprocessed, standardized and normalized, thereby improving the generalization ability of the model and avoiding the problem of gradient disappearance or explosion during the training process of the model.
[0098] In step S42, a convolution kernel of a preset size is used to perform at least one convolution operation and pooling operation on the feature vector to obtain a one-dimensional feature vector.
[0099] In the disclosed embodiment, the feature vector of the current frame is convolved through a convolution kernel, and the feature vector after the convolution operation is pooled to reduce the size of the feature vector after the convolution, further reducing the amount of model calculation. The convolution and pooling operations are repeated until a one-dimensional feature vector is obtained.
[0100] It should be understood that the convolution and pooling operations are performed alternately, and the number of convolution and pooling operations can be set according to actual needs.
[0101] In step S43, the one-dimensional feature vector is compared with the preset feature for similarity.
[0102] In the disclosed embodiment, the preset features are used to represent the feature vector after feature extraction of the image with the stuck pop-up window.
[0103] In the disclosed embodiment, a new vector can be obtained by subtracting a one-dimensional feature vector from a preset feature vector, and the new vector is input into a fully connected layer to obtain a scalar, which can be used to represent the degree of similarity between the one-dimensional feature vector after feature extraction of the image of the current frame and the preset feature vector.
[0104] In step S44, in response to the similarity being greater than the second threshold, it is determined that there is a freeze.
[0105] In the disclosed embodiment, a sigmoid function can be used to map a scalar representing the degree of similarity between a one-dimensional feature vector after feature extraction of an image of a current frame and a preset feature vector to 0 or 1. For example, assuming that the first threshold is 0.5, if the scalar is greater than or equal to 0.5, it means that the predicted value of the current frame is similar to the current frame, and a mapping value of 1 is obtained, indicating that there is a freeze; if the scalar is less than 0.5, it means that the predicted value of the current frame is not similar to the current frame, and a mapping value of 0 is obtained, indicating that there is no freeze.
[0106] In an exemplary embodiment, the second model is a six-layer convolutional neural network, and the second model consists of six convolutional layers and six pooling layers. Use a convolution kernel of size 3*3 to extract features of the current frame image, and finally obtain a one-dimensional feature vector corresponding to the current frame image, and use relu as the activation function and sigmoid function as the loss function as the fully connected layer for output, and compare the similarity of the final output one-dimensional vector features with the preset features. If the similarity is greater than 0.5, it means that a jamming pop-up window appears, and it is determined that there is a jamming phenomenon; if the similarity is less than 0.5, it means that no jamming pop-up window appears, and it is determined that there is no jamming phenomenon.
[0107] In the disclosed embodiment, the first model is trained by inputting training videos and test videos into the first model.
[0108] Figure 5 is a flowchart of a first model training method according to an exemplary embodiment. Figure 5 As shown, the following steps are included.
[0109] In step S51, a training video and a test video are obtained, and the training video and the test video are segmented to obtain i training frames and j test frames respectively.
[0110] In the disclosed embodiment, i and j are positive integers. The training video and the test video are segmented respectively, and one frame of image is obtained each time segmentation, and after segmentation, i training frames and j test frames are obtained.
[0111] In step S52, feature extraction is performed on the i training frames and the j test frames to obtain feature vectors corresponding to the i training frames and the j test frames, respectively.
[0112] In the disclosed embodiment, the feature vectors corresponding to the acquired i training frames can be sorted according to the time sequence in the training video to obtain a training sequence. The feature vectors corresponding to the acquired j test frames can be sorted according to the time sequence in the test video to obtain a training sequence, which is convenient for inputting into the first model for training and testing.
[0113] In step S53, the feature vectors of i training frames are subtracted to obtain the difference vectors of i-1 training frames.
[0114] In the disclosed embodiment, adjacent feature vectors in the feature vectors of i training frames are subtracted from each other to obtain difference vectors of i-1 training frames, and the difference vectors of i-1 training frames can reflect the changes between the training frames.
[0115] In step S54, the feature vectors of the j test frames are subtracted to obtain the difference vectors of j-1 test frames.
[0116] In the disclosed embodiment, adjacent feature vectors in the feature vectors of j test frames are subtracted from each other to obtain difference vectors of j-1 test frames, and the difference vectors of j-1 test frames can reflect the changes between the test frames.
[0117] In step S55, the difference vectors of the i-1 training frames are input into the first prediction model for training, and the first prediction model is constrained based on the difference vectors of the j-1 test frames to obtain the first model through training.
[0118] In the disclosed embodiment, the difference vectors of i-1 training frames are used for training, and the model parameters can be updated through a back-propagation algorithm and an optimizer, so that the model can better fit the training data.
[0119] In the disclosed embodiment, the performance of the first model is evaluated on the difference vectors of j-1 test frames, for example, accuracy, precision, recall, F1 score, etc.
[0120] In an exemplary embodiment, the first model is LSTM, a training video, a verification video and a defined LSTM are prepared, a time series consisting of feature vectors corresponding to 50 verification frames in the training video is obtained, and 49 difference vectors are obtained by subtracting the feature vectors corresponding to adjacent frames in the time series, and the 49 difference vectors are input into the LSTM, and the output of each layer is calculated by forward propagation, and then the loss function is calculated, and the parameters in the network are updated by back propagation. Through repeated iterations, until the number of iterations or the loss function value reaches a preset threshold, the training using the feature vector of the training frame is terminated. The LSTM is evaluated using the feature vector difference corresponding to 50 adjacent frames in the verification video. During the verification process, the data in the verification set is input, the output is calculated by forward propagation, and the evaluation index on the verification set is calculated to evaluate the performance of the model, and the parameters and model structure are adjusted based on the evaluation results. The final trained model is tested using the test set. In addition, the LSTM after training and verification can also be introduced in the test video for testing, the data in the test video is input, the output is calculated by forward propagation, and the evaluation index on the test set is calculated to evaluate the performance of the model.
[0121] Figure 6 is a flowchart of a second model training method according to an exemplary embodiment. Figure 6 As shown, the following steps are included.
[0122] In step S61, the feature vectors corresponding to the sample images are marked, the feature vectors corresponding to the images without pop-up windows or with non-stuttering pop-up windows are marked as positive samples, and the feature vectors of the images with stuttering pop-up windows are input into the second model as negative samples.
[0123] In the disclosed embodiment, by inputting the feature vector corresponding to the sample image containing the positive sample or negative sample label into the second model, the second model performs supervised learning, so that the second model can accurately classify the data.
[0124] In step S62, feature extraction is performed on the feature vector through multiple convolutional layers and pooling layers to obtain an extracted feature vector.
[0125] In the disclosed embodiment, the number of convolutional layers and pooling layers can be set based on demand and can also be adjusted based on test results.
[0126] In step S63, the model parameters in the second model are adjusted based on the extracted feature vector.
[0127] In the disclosed embodiment, based on the extracted feature vector, a back propagation algorithm and an optimizer are used to update the parameters in the network.
[0128] In step S64, the above process is repeated until the recognition result of the second model meets the preset effect, thereby obtaining the second model.
[0129] In the disclosed embodiment, the process of training the second model using the feature vector corresponding to the sample image is repeated for iteration. After each iteration, the accuracy of the second model is tested using the verification data, and in response to detecting that the accuracy of the second model meets the requirement, it is determined that the training of the second model is completed.
[0130] In an exemplary embodiment, the second model is a six-layer CNN, and the training data is a picture marked as a positive sample that does not contain a pop-up window or contains a non-stuttering pop-up window, and a picture marked as a negative sample that contains a stuttering pop-up window. The picture can be labeled by adding a label to the picture data. The picture data is subjected to preprocessing operations including preprocessing and standardization. The training data is input into the CNN model for forward propagation calculation. In each layer, the data undergoes operations such as convolution and pooling to finally obtain the output result. The output result of the model is compared with the true label to calculate the value of the loss function. Based on the information of the loss function, the parameters of the CNN model are updated in reverse to reduce the value of the loss function. Repeatedly perform forward propagation calculations, calculate the value of the loss function, and reversely update the parameters of the CNN model to achieve multiple iterative training. In response to meeting the number of training times, the iteration is stopped, and the accuracy and precision of the model are tested using test data. According to the test results, the model is optimized to finally complete the training process of the CNN.
[0131] In the disclosed embodiment, when the jam detection result obtained based on the network model is jam, it indicates that the current frame has jammed, and the terminal jam log or jam report recorded can be captured to facilitate maintenance and debugging. In addition, by obtaining the image information of the current frame where the jam occurs or the video information containing the current frame, it is convenient to intuitively view the jam phenomenon and debug, and it is also convenient to detect the correctness of the jam detection result. It should be understood that the format of the captured log or report, as well as the image information and video information, can be any format determined based on actual needs, and this is not limited in the disclosed embodiment.
[0132] Figure 7 The figure is a schematic diagram of a jamming detection scenario according to an exemplary embodiment.
[0133] In an exemplary embodiment, the first model is an LSTM model, and the second model is a multi-layer CNN model. Computer 1 automatically controls mobile phone 2 to operate a test application (APP). By using the multi-layer CNN model for image recognition and the LSTM model using the inverse difference algorithm to analyze the current frame, scoring results 1 and 2 are obtained, and then the two results are ORed to obtain result 3. If result 3 is True, the loop continues to detect whether the APP is stuck. If result 3 is False, the error log (Bug) is captured and the video is recorded to retain the current frame.
[0134] Figure 8 is a flowchart of training a first model according to an exemplary embodiment.
[0135] In an exemplary embodiment, the first model is an LSTM model. By importing a training video and a test video, the test video has as many scene switches as possible and is relatively complex. Then, the video set is divided to obtain a training set data1 and a test set data2 consisting of a frame difference array, and then the difference value of the adjacent frames of data1 and data2 is calculated, and the difference is stored in a new two-dimensional array. Use linear interpolation to scale data1 and data2 to the interval of [-1,1], and adjust the size of the data to meet the input requirements of the model. After scaling data1 and data2, use the data of data1 to train the LSTM, and use the data of data2 to test, and finally obtain a trained LSTM. The advantage of choosing the LSTM model for training is that it has long-term memory, can store associated information, and works well in processing serialized data.
[0136] Fig. 9 is a flowchart of using a first model according to an exemplary embodiment.
[0137] In an exemplary embodiment, the first model is an LSTM model. It should be understood that in this exemplary embodiment, the first sequence may also be referred to as an original sequence, and the second sequence may also be referred to as a new sequence. The original time series data is differentiated by using a differential method to differentiate two adjacent frames in the original sequence to obtain a new sequence, and the new sequence is input into the LSTM for prediction to obtain a predicted difference between a feature vector corresponding to the current frame and a feature vector corresponding to the previous frame of the current frame. Through a reverse differential method, based on the feature vector corresponding to the previous frame of the current frame and the predicted difference, a predicted value of the feature vector corresponding to the current frame is obtained, and finally the predicted value is tested to obtain the result of the jam detection.
[0138] Fig.10 is a schematic diagram showing the use of a second model structure according to an exemplary embodiment.
[0139] exist Fig.10In the figure, the second model is a CNN model, where the CNN model structure is 4 convolutional layers and 4 pooling layers. The Layer (type) column is used to indicate the name and type of each layer, the Output Shape column is used to indicate the size and number of channels of the graphic features output by each layer, and Param is used to indicate the parameter information of the input of each layer. The parameters of each convolutional layer can be different. As the number of layers increases, the size of the output graphic gradually decreases, and the number of channels gradually increases.
[0140] Fig.11 is a schematic diagram of a training process using a second model according to an exemplary embodiment.
[0141] exist Fig.11 In , when the second model is a CNN model, the input of the model can be set to a 200*200 three-channel image, the convolution kernel size is 3*3, relu is used as the activation function, the Sigmoid function is used for classification, and the variable is mapped to 0 or 1. Assuming that the number of negative samples used is about 150, the use of a large model will result in overfitting to a certain extent, and because the large model has a complex network, it affects the recognition accuracy, so a 6-layer convolutional neural network is selected. Using the exponentially weighted moving average (RMSprop) as the optimizer can reduce the number of iterations required to reach the optimal value and improve the ability of the optimization algorithm. The input of the CNN model is a picture, and the picture features are learned through the neural network. The output is the probability value of each category. The larger the probability value, the index category corresponding to the corresponding large probability value is the result value. Fig.11 In the example, the content between each Epoch is used to represent the process of each iteration. "Epoch 5 / 80" is used to indicate that the current iteration is the fifth iteration, and the total number of iterations is 80. The content above "Epoch 5 / 80" is used to indicate the training process of the CNN model. There are 12 lines in total, indicating that the training of the CNN is divided into 12 steps. After the completion of steps 1 to 11, the time consumption (ETA), loss (loss), and accuracy (acc) are output. After the completion of step 12, that is, after the completion of this iteration, the total time consumption of this iteration, the average time consumption of each step, the loss (loss) of step 12, the accuracy (acc), the average loss (val_loss) of this iteration, and the average accuracy (val_acc) are obtained.
[0142] Fig.12 is a schematic diagram showing the change in accuracy of a training process using a second model according to an exemplary embodiment.
[0143] Fig.12 In the figure, the horizontal axis represents the number of iterations, and the vertical axis represents the accuracy. As the number of iterations increases, it can be determined that the training accuracy and the verification accuracy continue to increase with the number of iterations. This shows that multiple iterations of the second model can improve the accuracy of the second model processing.
[0144] Fig.13 It is a schematic diagram showing the change of loss value in the training process using the second model according to an exemplary embodiment.
[0145] Fig.13 In the figure, the horizontal axis represents the number of iterations, and the vertical axis represents the loss value. As the number of iterations increases, it can be determined that the training accuracy and the verification accuracy continue to decrease with the increase in the number of iterations and eventually approach 0. This shows that multiple iterations of the second model training can reduce the loss value of the second model processing.
[0146] In an exemplary embodiment, during the process of training and tuning the model, some noise is injected into the positive samples to reduce the model's misjudgment rate. For example, some system pop-up notifications are normal phenomena. In order to prevent model recognition errors, we will add a certain amount of noise to the positive samples to make the model's generalization performance stronger.
[0147] Fig.14 It is a schematic diagram showing the change of training effect using the first model according to an exemplary embodiment.
[0148] Fig.14 In the figure, the horizontal axis represents the number of iterations of the first model training, and the vertical axis represents the feature information corresponding to the frame image. Through multiple training iterations of the first model, the feature gap between the predicted results of the first model and the actual results gradually decreases, indicating that the training effect of the first model is obvious.
[0149] In an exemplary embodiment, the following code may be designed to implement inverse difference.
[0150] definvent_scale(sef,scaler,X,y) : Defines a function called invent_scale that accepts four parameters: sef, scaler, X, and y.
[0151] new_row = [x for x in X] + [y]: uses list comprehension to collect each element in X (assumed to be a single number or data point) into the list new_row and adds y to the end of this list.
[0152] array = np.array(new_row) : Converts a list to an array.
[0153] array = array.reshape(1,len(array)): Reshape the array into a two-dimensional array with shape [1,2].
[0154] invert = scaler.inverse_transform (array): inverse scaling
[0155] return invert[0,-1]: Returns the inverse scaled y value
[0156] In an exemplary embodiment, the following code may be designed to implement the construction of an LSTM model.
[0157] deffit_lstm(self,train,batch_size,nb_epoch,neurons):: Defines a method named fit_lstm that accepts five parameters: self (usually omitted in classes), train, batch_size, nb_epoch, and neurons.
[0158] X,y = train[:,:-1],train[:,-1]: Use the slicing operation to separate the x and y in the training data.
[0159] X = X.reshape(x.shape[o], 1, x.shape[1]) : Use NumPy’s reshape function to convert two-dimensional data into three-dimensional data.
[0160] model = Sequential(): Initialize a Sequential model.
[0161] model.add(LSTM(neurons, batch_input_shape = (batch_size, X.shape[1], X.shape[2]), stateful = True)) : Adds an LSTM layer to the Sequential model and sets the parameters.
[0162] modet.add(Dense(1)) : Adds a fully connected layer with a single output neuron to the Sequential model.
[0163] model.compile(loss = 'mean_squared_error', optimizer = 'adam') : compiles the model using mean squared error as the loss function and the Adam optimizer.
[0164] for iin range(nb_epoch): Start a loop for training the model.
[0165] his = model.fit(X, y, batch_size = batch_size, verbose = 1, shuffle = False) : Trains the model without shuffling the data, outputs verbose information, and uses the given batch size.
[0166] model.reset_states(): reset the network state after each training. The network state is different from the network weight.
[0167] return model
[0168] Fig.15 is a schematic diagram of a normal pop-up window according to an exemplary embodiment.
[0169] exist Fig.15 In the pop-up window, there is a pop-up window, but the pop-up window content does not contain content indicating that there is a jam and no response, so it is not a jam pop-up window, so you can Fig.15 The positive sample training data of the second model is input into the second model.
[0170] Fig.16 The figure is a schematic diagram of a stuck pop-up window according to an exemplary embodiment.
[0171] exist Fig.16 In the pop-up window, the content is "XXX is not responding", so it is an image containing an unresponsive pop-up window caused by lag, indicating that there is a lag phenomenon, so you can Fig.16 The negative sample training data is input into the second model.
[0172] In the disclosed embodiment, a network model is used to detect whether there is a freeze phenomenon in the current frame, thereby accurately detecting freeze pop-up windows and freeze conditions during terminal use, thereby improving the freeze detection range, and automatically capturing logs or recording videos based on the freeze information, saving human resources.
[0173] Based on the same concept, an embodiment of the present disclosure also provides a freeze detection device.
[0174] It is understandable that the jam detection device provided by the embodiment of the present disclosure includes hardware structures and / or software modules corresponding to the execution of each function in order to realize the above functions. In combination with the units and algorithm steps of each example disclosed in the embodiment of the present disclosure, the embodiment of the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art may use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the technical solution of the embodiment of the present disclosure.
[0175] Fig.17 is a block diagram of a jamming detection device 100 according to an exemplary embodiment. Fig.17 The device includes a collection unit 101 and a processing unit 102.
[0176] The acquisition unit 101 is used to acquire n frames of images including the current frame in the terminal display screen, where n is a positive integer greater than or equal to 2;
[0177] The processing unit 102 is used to call the network model for jamming detection, and obtain a jamming detection result based on the network model and n frames of images; wherein the jamming detection result includes the presence of jamming or the absence of jamming.
[0178] In one embodiment, the network model includes a first model and a second model, the first model is used to obtain a jam detection result by performing a similarity test on n frames of images, and the second model is used to obtain a jam detection result by performing a feature extraction on the current frame and performing a similarity test between the processed features and the preset features; the processing unit obtains the jam detection result based on the network model and the n frames of images in the following manner: in response to the first model or the second model determining that jams occur in the n frames of images, the jam detection result is determined to be the presence of jams.
[0179] In one embodiment, the network model includes a first model; the processing unit 102 obtains a freeze detection result based on the network model and n frames of images in the following manner:
[0180] Acquire feature vectors corresponding to n frames of images, and arrange the feature vectors corresponding to the n frames of images in chronological order to obtain a first sequence; subtract feature vectors of two adjacent frames in the first sequence to obtain a second sequence containing n-1 difference vectors; input the second sequence into the first model to obtain the first model's predicted value of the difference vector between the current frame and the frame before the current frame; determine the predicted value of the current frame based on the predicted value; compare the predicted value of the current frame with the current frame for similarity; in response to the similarity being less than a first threshold, determine that a freeze exists.
[0181] In one embodiment, the network model includes a second model; the processing unit 102 obtains a freeze detection result based on the network model and n frames of images in the following manner:
[0182] Obtain a feature vector of the current frame; perform at least one convolution operation and pooling operation on the feature vector using a convolution kernel of a preset size to obtain a one-dimensional feature vector; compare the one-dimensional feature vector with the preset feature for similarity; and determine that a freeze exists in response to the similarity being greater than a second threshold.
[0183] In one embodiment, the first model is trained in the following manner: a training video and a test video are obtained, and the training video and the test video are respectively segmented to obtain i training frames and j test frames, where i and j are positive integers; feature extraction is performed on the i training frames and the j test frames to obtain feature vectors corresponding to the i training frames and the j test frames, respectively; the feature vectors of the i training frames are subtracted to obtain difference vectors of the i-1 training frames; the feature vectors of the j test frames are subtracted to obtain difference vectors of the j-1 test frames; the difference vectors of the i-1 training frames are input into the first prediction model for training, and the first prediction model is constrained based on the difference vectors of the j-1 test frames to obtain the first model through training.
[0184] In one embodiment, the second model is trained in the following manner:
[0185] The sample images are marked, and the images without pop-up windows or with non-stuttering pop-up windows are marked as positive samples, and the images with stuttering pop-up windows are input into the second model as negative samples; the features of the sample images are extracted through multiple convolutional layers and pooling layers to obtain feature maps; based on the feature maps, the model parameters in the second model are adjusted; the above process is repeated until the recognition result of the second model meets the preset effect, and the second model is obtained.
[0186] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0187] Fig.18 2 is a block diagram of a device 200 for jamming detection according to an exemplary embodiment. For example, the device 200 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0188] Reference Fig.18 , the device 200 may include one or more of the following components: a processing component 202 , a memory 204 , a power component 206 , a multimedia component 208 , an audio component 210 , an input / output (I / O) interface 212 , a sensor component 214 , and a communication component 216 .
[0189] The processing component 202 generally controls the overall operation of the device 200, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 202 may include one or more processors 220 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 202 may include one or more modules to facilitate interaction between the processing component 202 and other components. For example, the processing component 202 may include a multimedia module to facilitate interaction between the multimedia component 208 and the processing component 202.
[0190] The memory 204 is configured to store various types of data to support operations on the device 200. Examples of such data include instructions for any application or method operating on the device 200, contact data, phone book data, messages, pictures, videos, etc. The memory 204 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0191] The power component 206 provides power to the various components of the device 200. The power component 206 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 200.
[0192] The multimedia component 208 includes a screen that provides an output interface between the device 200 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 208 includes a front camera and / or a rear camera. When the device 200 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.
[0193] The audio component 210 is configured to output and / or input audio signals. For example, the audio component 210 includes a microphone (MIC), and when the device 200 is in an operation mode, such as a call mode, a recording mode, and a speech recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 204 or sent via the communication component 216. In some embodiments, the audio component 210 also includes a speaker for outputting audio signals.
[0194] I / O interface 212 provides an interface between processing component 202 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a start button, and a lock button.
[0195] The sensor assembly 214 includes one or more sensors for providing various aspects of the status assessment of the device 200. For example, the sensor assembly 214 can detect the open / closed state of the device 200, the relative positioning of components, such as the display and keypad of the device 200, the sensor assembly 214 can also detect the position change of the device 200 or a component of the device 200, the presence or absence of user contact with the device 200, the orientation or acceleration / deceleration of the device 200 and the temperature change of the device 200. The sensor assembly 214 can include a proximity sensor configured to detect the presence of a nearby object without any physical contact. The sensor assembly 214 can also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 214 can also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor or a temperature sensor.
[0196] The communication component 216 is configured to facilitate wired or wireless communication between the device 200 and other devices. The device 200 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 216 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 216 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0197] In an exemplary embodiment, the apparatus 200 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components to perform the above method.
[0198] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 204 including instructions, and the instructions can be executed by the processor 220 of the device 200 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0199] It is to be understood that in the present disclosure, "plurality" refers to two or more than two, and other quantifiers are similar. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. The singular forms "a", "the" and "the" are also intended to include plural forms, unless the context clearly indicates other meanings.
[0200] It is further understood that the terms "first", "second", etc. are used to describe various information, but such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other, and do not indicate a specific order or degree of importance. In fact, the expressions "first", "second", etc. can be used interchangeably. For example, without departing from the scope of the present disclosure, the first information can also be referred to as the second information, and similarly, the second information can also be referred to as the first information.
[0201] It can be further understood that, unless otherwise specified, “connection” includes a direct connection without other components between the two, and also includes an indirect connection with other components between the two.
[0202] It is further understood that, although the operations are described in a specific order in the drawings in the embodiments of the present disclosure, it should not be understood as requiring the operations to be performed in the specific order shown or in a serial order, or requiring the execution of all the operations shown to obtain the desired results. In certain environments, multitasking and parallel processing may be advantageous.
[0203] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any modifications, uses or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common knowledge or customary technical means in the art that are not disclosed in the present disclosure.
[0204] It should be understood that the present disclosure is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the scope of the appended claims.
Claims
1. A jam detection method, characterized in that: include: Acquire n frames of images, wherein the n frames of images include a current frame of a terminal display screen, and n is a positive integer greater than 2; Calling a network model for jamming detection, and obtaining a jamming detection result based on the network model and the n frames of images; The jam detection result includes the presence of jam or the absence of jam.
2. The method according to claim 1, characterized in that The network model includes a first model and a second model, the first model is used to obtain a jamming detection result by performing similarity detection on the n frames of images, and the second model is used to extract features from the current frame and perform similarity detection between the processed features and the preset features to obtain a jamming detection result; The obtaining of a jamming detection result based on the network model and the n frames of images includes: In response to the first model or the second model determining that the n frames of images have jammed, the jam detection result is determined to be the presence of jamming.
3. The method according to claim 2, characterized in that The network model includes a first model; The obtaining of a jamming detection result based on the network model and the n frames of images includes: Acquire feature vectors corresponding to n frames of images, and arrange feature vectors corresponding to n-1 frames of images in chronological order to obtain a first sequence, wherein the n-1 frames of images do not include a current frame of a terminal display screen; Subtract the feature vectors of two adjacent frames in the first sequence to obtain a second sequence containing n-2 difference vectors; Inputting the second sequence into the first model, obtaining a prediction value of the first model for a difference vector between a current frame and a frame before the current frame; Based on the predicted value, determining a predicted value of the current frame; Comparing the predicted value of the current frame with the current frame for similarity; In response to the similarity being less than a first threshold, it is determined that a freeze exists.
4. The method according to claim 2, characterized in that: The network model includes a second model; The obtaining of a jamming detection result based on the network model and the n frames of images includes: Get the feature vector of the current frame; Using a convolution kernel of a preset size to perform at least one convolution operation and a pooling operation on the feature vector to obtain a one-dimensional feature vector; Comparing the one-dimensional feature vector with the preset feature for similarity; In response to the similarity being greater than a second threshold, it is determined that a freeze exists.
5. The method according to claim 2 or 3, characterized in that: The first model is trained in the following way: Obtain a training video and a verification video, and segment the training video and the verification video to obtain i training frames and j verification frames, respectively, where i and j are positive integers; Perform feature extraction on the i training frames and the j verification frames to obtain feature vectors corresponding to the i training frames and the j verification frames respectively; Subtract the feature vectors of two adjacent frames of the i training frames to obtain a difference vector of i-1 training frames; Subtract the feature vectors of two adjacent frames of the j verification frames to obtain the difference vectors of j-1 verification frames; The difference vectors of the i-1 training frames are input into a first prediction model for training, and the first prediction model is constrained based on the difference vectors of the j-1 verification frames to obtain a first model through training.
6. The method according to claim 2 or 4, characterized in that: The second model is trained in the following way: Mark the feature vectors corresponding to the sample images, take images without pop-up windows or with non-stuttering pop-up windows as positive samples, and input images with stuttering pop-up windows into the second model as negative samples; Perform feature extraction on the image through multiple convolutional layers and pooling layers to obtain an extracted feature vector; Based on the extracted feature vector, adjusting model parameters in the second model; Repeat the above process until the recognition result of the second model meets the preset effect, thereby obtaining the second model.
7. The method according to claim 1, characterized in that The method further comprises: In response to the jam detection result indicating that jamming exists, at least one of the following operations is performed: Capture logs containing freeze information from the terminal; Saving the image information of the current frame; The video information including the current frame is saved.
8. A jamming detection device, characterized in that: include: A collection unit, used to acquire n frames of images, wherein the n frames of images include a current frame of a terminal display screen, and n is a positive integer greater than or equal to 2; A processing unit, used to call a network model for jamming detection, and obtain a jamming detection result based on the network model and the n frames of images; The jam detection result includes the presence of jam or the absence of jam.
9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to: execute the method described in any one of claims 1 to claim 7.
10. A storage medium, characterized in that: The storage medium stores instructions, and when the instructions in the storage medium are executed by a processor of the terminal, the terminal is enabled to execute the method described in any one of claims 1 to 7.