Labeling method, broken screen recognition model construction method and broken screen recognition method
By marking the blocked screen extension and core area marking boxes on the screen image, and the sampling probability distribution calculation and learning marking boxes are solved, the low accuracy and overfitting of the broken screen recognition model are solved, and efficient broken screen recognition is achieved.
Patent Information
- Application Number
- CN202311865254.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-01
AI Technical Summary
In the prior art, the range of broken screen labeling frames is too large, resulting in slow convergence of the learning process of deep learning neural networks, and low accuracy of broken screen recognition, and there are problems of false detection and missed detection.
By obtaining the first and second annotation boxes including the broken screen extension area and the core area of the screen, the probability distribution of the initial learning annotation box is obtained by sampling the distance relationship between the borders, and the final learning annotation box is calculated and marked on the screen image to enhance the accuracy of data annotation.
It improves the accuracy of the screen break recognition model, reduces errors, avoids the risk of overfitting of neural networks, and achieves fast and accurate screen break recognition.
Smart Images

Figure CN120236158A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a labeling method, a method for constructing a broken screen recognition model, and a broken screen recognition method. Background Art
[0002] With the development and improvement of smart phone technology, social activities and work tasks are inseparable from mobile phones. However, during daily use, due to reasons such as accidental drops and physical impacts, the mobile phone screen may be broken. During the repair process of damaged mobile phones, both express delivery companies and repair manufacturers will take pictures of the mobile phone screen for record to clarify the degree of damage and the attribution of responsibility.
[0003] Currently, deep learning neural networks are mainly used to detect broken screens. Due to the large differences in the crack density, extension length, and extension direction of broken screens, simple rectangular bounding box labeling may include a large area of irrelevant regions, resulting in a slow convergence of the learning process of deep learning neural networks, and there is also a risk of overfitting. In addition, the neural network after deep learning is prone to problems such as false detection and missed detection of broken screens. Summary of the Invention
[0004] The present invention provides a labeling method, a method for constructing a broken screen recognition model, and a broken screen recognition method to solve the technical problem in the prior art that the range of the broken screen bounding box is too large, which easily leads to a slow convergence of the learning process of deep learning neural networks, and further results in low accuracy and large errors in the network model for detecting broken screens.
[0005] To solve the above technical problem, an embodiment of the present invention provides a labeling method, including:
[0006] Obtaining a screen image labeled with a first bounding box and a second bounding box; wherein, the first bounding box includes the broken screen extension area of the screen, and the second bounding box includes the broken screen core area of the screen;
[0007] Sampling to obtain the probability distribution of the initial learning bounding box according to the distance relationship between the borders of the first bounding box and the second bounding box;
[0008] Calculating to obtain the final learning bounding box according to the probability distribution of the initial learning bounding box, and labeling it on the screen image.
[0009] As a preferred solution, the sampling to obtain the probability distribution of the initial learning bounding box according to the distance relationship between the borders of the first bounding box and the second bounding box includes:
[0010] Obtaining the number of rows and columns of the pixel values occupied by a horizontal side and a vertical side in the second bounding box in the screen image respectively, and taking the number of rows as the first mean value and the number of columns as the second mean value;
[0011] Calculate the distances that all sides of the second annotation box are translated and extended to the corresponding sides of the first annotation box according to the positions of the first annotation box and the second annotation box in the screen image; wherein, the area of the screen image covered by the second annotation box is within the area of the screen image covered by the first annotation box, and each side of the second annotation box corresponds to one side of the first annotation box;
[0012] Scale the extended distances of all sides according to a preset ratio as the variance corresponding to that side; wherein, each side corresponds to a variance;
[0013] Generate the probability distribution of each side of the initial learning annotation box according to the first mean, the second mean, and the variance corresponding to each side.
[0014] As a preferred solution,
[0015] The initial learning annotation box includes a first side, a second side, a third side, and a fourth side; the first side and the second side are respectively the horizontal side and the vertical side forming an angle in the initial learning annotation box, and the third side and the fourth side are respectively the horizontal side and the vertical side forming the other diagonal of the initial learning annotation box;
[0016] The length of the first side is greater than the first mean, the length of the second side is greater than the second mean, the length of the third side is less than the first mean, and the length of the fourth side is less than the second mean.
[0017] As a preferred solution, calculating the final learning annotation box according to the probability distribution of the initial learning annotation box and annotating it on the screen image includes:
[0018] Calculate the position probability, length probability, and width probability of the center point of the final learning annotation box according to the probability distribution of each side of the initial learning annotation box;
[0019] Obtain the final learning annotation box according to the position probability, length probability, and width probability of the center point, and annotate the final learning annotation box on the screen image.
[0020] Correspondingly, the present invention also provides a method for constructing a broken screen recognition model, the method including:
[0021] Obtain training samples; wherein, the training samples include at least one original screen image and at least one annotated image annotated with a learning annotation box corresponding to the original screen image; the annotated image annotated with the learning annotation box is obtained by the annotation method described in any one of the above;
[0022] Construct a preset neural network, and input the training samples into the neural network for training to obtain a cracked screen recognition model.
[0023] As a preferred solution, the step of constructing a preset neural network and inputting the training samples into the neural network for training to obtain a cracked screen recognition model is specifically as follows:
[0024] Construct a preset neural network;
[0025] According to the labeled images with learning labeled boxes in the training samples, obtain the center point position, length, and width information of the learning labeled boxes;
[0026] According to the center point position, length, and width information of the learning labeled boxes in each labeled image, calculate the loss function corresponding to each labeled image;
[0027] Input the training samples into the neural network, and combine the loss function corresponding to each labeled image to train and obtain a cracked screen recognition model.
[0028] Correspondingly, the present invention also provides a cracked screen recognition method, and the method includes:
[0029] Obtain an image of the screen to be recognized;
[0030] Input the image of the screen to be recognized into the cracked screen recognition model, so that the cracked screen recognition model performs cracked screen recognition on the image of the screen to be recognized; the cracked screen recognition model is obtained by the construction method of the cracked screen recognition model described in any one of the above;
[0031] When a cracked screen is recognized, generate a labeled box labeled with the cracked screen area on the image of the screen to be recognized.
[0032] Correspondingly, the present invention also provides a labeling device, including: a screen image module, a sampling module, and a labeling module;
[0033] The screen image module is used to obtain a screen image labeled with a first labeled box and a second labeled box; wherein, the first labeled box includes the cracked screen extension area of the screen, and the second labeled box includes the cracked screen core area of the screen;
[0034] The sampling module is used to sample the probability distribution of the initial learning labeled box according to the distance relationship between the first labeled box and the second labeled box;
[0035] The labeling module is used to calculate the final learning labeled box according to the probability distribution of the initial learning labeled box and label it on the screen image.
[0036] Correspondingly, the present invention also provides a device for constructing a broken screen recognition model, including: a sample module and a training module;
[0037] The sample module is used to obtain training samples; wherein, the training samples include at least one original screen image and at least one annotated image with a learned annotation box corresponding to the original screen image;
[0038] The training module is used to construct a preset neural network and input the training samples into the neural network for training, so as to obtain a broken screen recognition model;
[0039] Among them, the annotated image of the training sample in the method for constructing the broken screen recognition model is obtained by the annotation method described in any one of the above.
[0040] Correspondingly, the present invention also provides a broken screen recognition device, including: a screen image module, a recognition module and an output module;
[0041] The screen image module is used to obtain a screen image to be recognized;
[0042] The recognition module is used to input the screen image to be recognized into the broken screen recognition model, so that the broken screen recognition model performs broken screen recognition on the screen image to be recognized;
[0043] The output module is used to generate an annotation box with a broken screen area annotated on the screen image to be recognized when a broken screen is recognized;
[0044] Among them, the broken screen recognition model is constructed by the method for constructing a broken screen recognition model described in any one of the above, and the annotated image of the training sample in the method for constructing the broken screen recognition model is obtained by the annotation method described in any one of the above.
[0045] Correspondingly, the present invention also provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the annotation method described in any one of the above, the method for constructing a broken screen recognition model described in any one of the above, or the broken screen recognition method described above.
[0046] Correspondingly, the present invention also provides a computer-readable storage medium, which includes a stored computer program. Among them, when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the annotation method described in any one of the above, the method for constructing a broken screen recognition model described in any one of the above, or the broken screen recognition method described above.
[0047] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0048] The technical solution of the present invention enables quick positioning to the cracked screen extension area and the cracked screen core area through a screen image marked with the distance relationship between the first marking frame and the second marking frame, samples the probability distribution of the initial learning marking frame, and then calculates the final learning marking frame, which is marked on the screen image. The markings of the cracked screen extension area and the cracked screen core area can sample and calculate the true value of the learning marking frame, thereby realizing data augmentation of the marking frame, making the sampling result slightly perturbed within a certain range to improve the accuracy of the marked data and avoiding the problem of large errors in the learning marking frame used for model training.
[0049] Furthermore, through the learning marking frame after data augmentation, it is possible to achieve the situation where the result does not converge when the cracked screen recognition model is trained on the screen image marked with the learning marking frame, and at the same time, it can effectively reduce the overfitting risk of the neural network model brought by the traditional rectangular marking frame with too large a range and inaccurate marking. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 : Schematic diagrams of the radial cracked screen and the spider web cracked screen after the existing terminal screen is cracked;
[0051] Figure 2 : Schematic diagrams of the existing marking frame marking the radial cracked screen and the spider web cracked screen respectively;
[0052] Figure 3 : Flowchart of the steps of a marking method provided by an embodiment of the present invention;
[0053] Figure 4 : Schematic diagrams of the marking frame marking the radial cracked screen and the spider web cracked screen respectively provided by an embodiment of the present invention;
[0054] Figure 5 : Schematic diagram of the probability distribution of the learning marking frame provided by an embodiment of the present invention;
[0055] Figure 6 : Flowchart of the steps of a method for constructing a cracked screen recognition model provided by an embodiment of the present invention;
[0056] Figure 7 : Flowchart of the steps of a cracked screen recognition method provided by an embodiment of the present invention;
[0057] Figure 8 : Schematic diagram of the structure of a marking device provided by an embodiment of the present invention;
[0058] Figure 9 : Schematic diagram of the structure of a device for constructing a cracked screen recognition model provided by an embodiment of the present invention;
[0059] Figure 10 : A structural schematic diagram of a broken screen recognition device provided by an embodiment of the present invention. Specific embodiments
[0060] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0061] At present, during the use of terminal electronic devices, due to reasons such as accidental dropping and physical impact, the screen is prone to breakage. And the screen breakage mainly includes two breakage forms. The first is a radial broken screen with a main central break point and extending outward for breakage. The second is a cobweb-like broken screen without a central break point but with uniform dense cracks in a certain area, as shown in Figure 1 shown.
[0062] Furthermore, regarding the first broken screen form, since it has a central break point and broken screen cracks extending outward, the broken screen extension lines of the radial broken screen are thin and long. And currently, the annotation box is mainly a rectangular annotation box, so that the area covered by the annotation box when annotating the radial broken screen is relatively large, including a large area of the screen that is not damaged, as shown in Figure 2 shown, which makes the training of the neural network model for broken screen recognition using this image data prone to slow convergence during the learning process. At the same time, when performing broken screen recognition through this neural network model for broken screen recognition, problems such as false detection and missed detection are likely to occur.
[0063] Embodiment 1
[0064] Please refer to Figure 3 , a labeling method provided by an embodiment of the present invention, including the following steps S101 - S103:
[0065] Step S101: Obtain a screen image labeled with a first annotation box and a second annotation box; wherein, the first annotation box includes the broken screen extension area of the screen, and the second annotation box includes the broken screen core area of the screen.
[0066] In this embodiment, the labeling method is mainly used for labeling the broken screen area of the screen of a terminal device, especially for terminal electronic devices with display screens such as mobile phones, tablet computers, computers, and monitors.
[0067] In this embodiment, in order to avoid the anisotropic characteristics of the density, extension length, and extension direction of the cracked screen cracks, simple rectangular annotation frames may include a large area of irrelevant regions. Therefore, in the embodiment of the present invention, by obtaining a screen image including a first annotation frame of the cracked screen extension region of the screen and a second annotation frame of the cracked screen core region of the screen, the training sample data meeting the requirements of model training can be generated through the annotation frame data enhancement method of the following steps, so as to improve the accuracy and efficiency of cracked screen recognition of the generated cracked screen recognition model and avoid situations such as false detection and missed detection.
[0068] As an embodiment, the cracked screen core region may be the region with the largest crack density in the screen, that is, through the screen image marked with the distance relationship between the first annotation frame and the second annotation frame, the cracked screen extension region and the region with the largest crack density can be quickly located, thereby ensuring the accuracy of the learning annotation frame for crack range annotation.
[0069] In this embodiment, please refer to Figure 4 , the first annotation frame can be annotated by a conventional annotation method, that is, the first annotation frame completely surrounds a partial region of the cracked screen. The second annotation frame is annotated for the region with the largest crack density in the screen image; that is, for a radial cracked screen, the region annotated by the second annotation frame is the central break point (radial center) and a part within a preset range nearby; for a cobweb cracked screen, the region annotated by the second annotation frame is the point with the largest crack density in the cobweb and a part within a preset range nearby. Preferably, the preset range can be set according to actual model training accuracy, recognition rate, and other situations.
[0070] Step S102: Sample the probability distribution of the initial learning annotation frame according to the distance relationship between the first annotation frame and the second annotation frame.
[0071] As an embodiment, the sampling of the probability distribution of the initial learning annotation frame according to the distance relationship between the frames of the first annotation frame and the second annotation frame includes:
[0072] Obtain the number of rows and columns of pixels occupied by a horizontal side and a vertical side of the second bounding box in the screen image respectively, and use the number of rows as the first mean value and the number of columns as the second mean value; According to the positions of the first bounding box and the second bounding box in the screen image, calculate the distances that all sides of the second bounding box are translated and extended to the corresponding sides of the first bounding box; wherein, the screen image area covered by the second bounding box is within the screen image area covered by the first bounding box, and each side of the second bounding box corresponds to one side of the first bounding box; Scale the extended distances of all sides according to a preset ratio as the variance corresponding to each side; wherein, each side corresponds to a variance; Generate the probability distribution of each side of the initial learning bounding box according to the first mean value, the second mean value and the variance corresponding to each side.
[0073] In this embodiment, the number of rows or columns can be obtained by calculating the distances between the boundary coordinates of the second bounding box and the boundaries of the screen image. Specifically, the number of rows of pixels occupied by any horizontal side of the second bounding box in the screen image can be obtained, and the number of columns of pixels occupied by any vertical side of the second bounding box in the screen image can be obtained, and the number of rows and the number of columns are respectively used as the first mean value corresponding to the horizontal side and the second mean value corresponding to the vertical side. Furthermore, according to the positions of the first bounding box and the second bounding box in the screen image, calculate the distances that all sides of the second bounding box are translated and extended to the corresponding sides of the first bounding box. Among them, geometric transformation methods, such as affine transformation or perspective transformation, can be used to calculate the relative position relationship between the two bounding boxes.
[0074] In this embodiment, by traversing each side of the second bounding box, calculate the values of the extended distances of each side under the preset ratio, and use the values calculated for each side as the variance corresponding to each side. Furthermore, generate the probability distribution of each side of the initial learning bounding box through the mean value and the variance corresponding to each side. Among them, the Gaussian function can be used to describe the position information of each side, and the mean value represents the center of the position, and the variance represents the uncertainty of the position. As an embodiment, the probability distribution of each side of the initial learning bounding box can be a continuous probability distribution such as Gaussian probability distribution or chi-square distribution.
[0075] As an embodiment, the initial learning bounding box includes a first side, a second side, a third side and a fourth side; the first side and the second side are respectively the horizontal side and the vertical side that form an angle in the initial learning bounding box, and the third side and the fourth side are respectively the horizontal side and the vertical side that form the other diagonal angle in the initial learning bounding box;
[0076] The length of the first side is greater than the first mean value, the length of the second side is greater than the second mean value, the length of the third side is less than the first mean value, and the length of the fourth side is less than the second mean value.
[0077] In this embodiment, when generating the learning annotation boxes required for model training, the number of rows and columns of the pixel values occupied by the horizontal and vertical sides of the second annotation box in the screen image are used as the first mean M1 and the second mean M2. At the same time, the distances by which the second annotation box is translated and extended towards the corresponding sides of the first annotation box are calculated respectively. It can be understood that both the first annotation box and the second annotation box are rectangular boxes, so the number of sides of the first annotation box and the second annotation box is the same, such that for each side of the first annotation box in the same direction, there is a side of the second annotation box in the same direction, that is, these two sides are the corresponding sides of the first annotation box and the second annotation box in this direction.
[0078] Exemplarily, define the upper part of the screen image as the north direction, the lower part as the south direction, the left part as the west direction, and the right part as the east direction. Then, the distances by which the sides of the second annotation box in the four directions of east, west, south, and north are extended outwards to the corresponding four sides of the first annotation box are Le, Lw, Ls, and Ln respectively; where e represents the east direction, w represents the west direction, s represents the south direction, and n represents the north direction; the values of the preset ratios of Le, Lw, Ls, and Ln are calculated respectively as the variances corresponding to the sides, and then the probability distributions corresponding to the sides are generated. Preferably, the preset ratio is 1 / 3, so the generated Gaussian distribution is N(M, L / 3), where L belongs to [Le, Lw, Ls, Ln], M belongs to [M1, M2], and sampling is performed on the Gaussian distributions corresponding to each side:
[0079] X ∼ N + (M, L / 3), where X is the side of the initial learning annotation box towards the east or south, and X > M;
[0080] X ∼ N - (M, L / 3), where X is the side of the initial learning annotation box towards the west or north, and X > M.
[0081] Exemplarily, the sides of the initial learning annotation box towards the east and west are vertical sides, and the sides of the initial learning annotation box towards the south and north are horizontal sides. Furthermore, the distribution generated by the generated learning annotation box is as Figure 5 shown. Thus, according to the probability distributions corresponding to the sides of each initial learning annotation box, the probability of the center point position, the length probability, and the width probability of the final learning annotation box can be sampled and calculated. Among them, μ is the mean of the center, σ is the corresponding standard deviation, and the standard deviation σ can be calculated through the variance.
[0082] In this embodiment, by generating a probability distribution to describe the position information of the initial learning annotation box, the position of the initial learning annotation box in the screen image and the size of each side can be represented more accurately, which helps the model better learn the features of the object, thereby improving the accuracy of broken screen target detection. At the same time, using a probability distribution to represent the position information of the annotation box can better adapt to broken screen recognition targets of different shapes and sizes, enhancing the generalization ability of the learning model.
[0083] It can be understood that, compared with the simple bounding box representation method, the probability distribution can better describe the uncertainty of the object, so that the model has stronger generalization ability. At the same time, by calculating the variance of the extension distance, the relative position change of the learning annotation box between the first annotation box and the second annotation box can be considered, which helps the model learn the correlation between objects, further improving the performance of broken screen detection, and then realizing the improvement of model robustness. Using the mean and variance to generate the probability distribution can better handle the position error and uncertainty of the learning annotation box, improving the robustness of the model to noise and interference, so that the trained broken screen detection model can still maintain a high broken screen recognition performance in complex scenarios.
[0084] Step S103: Calculate the final learning annotation box according to the probability distribution of the initial learning annotation box, and mark it on the screen image.
[0085] As an embodiment, the calculating the final learning annotation box according to the probability distribution of the initial learning annotation box and marking it on the screen image includes:
[0086] Calculate the central point position probability, length probability and width probability of the final learning annotation box according to the probability distribution of each side of the initial learning annotation box; obtain the final learning annotation box according to the central point position probability, length probability and width probability, and mark the final learning annotation box on the screen image.
[0087] In this embodiment, through the probability distribution of each side of the initial learning annotation box, the probability distribution information about the position and size of the initial learning annotation box can be provided. Furthermore, through the probability distribution information of the position and size, the central point position probability, length probability and width probability of the final learning annotation box can be determined, so that the sampling result of the final learning annotation box is likely to be within a small range of perturbation within 1 standard deviation, avoiding the problem that the final learning annotation box does not converge or is difficult to converge, and effectively reducing the overfitting risk brought by traditional simple rectangular box annotation, so that the embodiment of the present invention can sample the true value of the broken screen area annotation box through the probability distribution method and use it as the learning target of the neural network to improve the accuracy of neural network training and reduce the error of broken screen recognition.
[0088] Implementing the above embodiments has the following effects:
[0089] The technical solution of the present invention enables quick positioning to the broken screen extension area and the broken screen core area through a screen image marked with the distance relationship between the first marking frame and the second marking frame, samples the probability distribution of the initial learning marking frame, and then calculates the final learning marking frame, which is marked on the screen image. The markings of the broken screen extension area and the broken screen core area can sample and calculate the true value of the learning marking frame, thereby realizing data augmentation of the marking frame, making the sampling results slightly perturbed within a certain range to improve the accuracy of the marked data, and avoiding the problem of large errors in the learning marking frame used for model training.
[0090] Embodiment 2
[0091] Please refer to Figure 6 , which is a method for constructing a broken screen recognition model provided by the present invention, implemented by the marking method described in any one of the above Embodiment 1, and includes the following steps S201 - S202:
[0092] Step S201: Obtain training samples; wherein, the training samples include at least one original screen image and at least one marked image marked with a learning marking frame corresponding to the original screen image; the marked image marked with the learning marking frame is obtained by the marking method described in any one of Embodiment 1.
[0093] In this embodiment, at least one set of training samples including the original screen image and the corresponding marked image marked with the learning marking frame is collected. Among them, the training samples are a data set for training a neural network model. The original screen image is an unprocessed broken screen image, and the marked image is an image marked with the broken screen area.
[0094] Step S202: Construct a preset neural network, and input the training samples into the neural network for training to obtain a broken screen recognition model.
[0095] In this embodiment, according to the task requirements and data characteristics, a preset neural network model is designed and constructed. Among them, the neural network model is a computational model of the structure and function of neurons. It is composed of multiple layers of neurons. Each neuron receives the input from the neurons in the previous layer, and after non - linear transformation through an activation function, outputs to the neurons in the next layer. The neural network model can learn to extract features from the input data and perform tasks such as classification or regression.
[0096] In this embodiment, the neural network model may include multiple convolutional layers, pooling layers, fully connected layers, etc., to achieve the recognition of cracked screen images. Thus, the collected training samples are input into the constructed neural network for training. During the training process, the neural network calculates the prediction results through forward propagation and updates the network parameters through the backpropagation algorithm to minimize the difference between the prediction results and the true labels. After multiple iterations of training, the neural network will gradually learn the feature representation of cracked screen images and be able to accurately identify the cracked screen area. The finally learned and trained neural network model is the cracked screen recognition model.
[0097] As an embodiment, constructing a preset neural network and inputting the training samples into the neural network for training to obtain a cracked screen recognition model specifically includes:
[0098] Construct a preset neural network; obtain the center point position, length, and width information of the learning annotation box according to the annotation image with a learning annotation box marked in the training sample; calculate the loss function corresponding to each annotation image according to the center point position, length, and width information of the learning annotation box of each annotation image; input the training samples into the neural network and combine the loss function corresponding to each annotation image to train and obtain a cracked screen recognition model.
[0099] In this embodiment, by constructing a preset neural network, the annotation image with a learning annotation box can be trained to improve the accuracy of cracked screen recognition. And according to the center point position, length, and width information of the learning annotation box of each annotation image, the loss function corresponding to each annotation image can be calculated. So that after the training samples are input into the neural network, the learning and training of the model can be optimized through the calculated loss function. Furthermore, after combining the loss function corresponding to each annotation image, the accuracy and precision of the model's cracked screen recognition can be optimized, thereby further reducing the error of model recognition. Further, in the embodiment of the present invention, by inputting the training samples into the neural network for training, an automated training process can be realized, reducing the need for manual intervention, enabling full-process automatic learning and training, and improving the efficiency of model training.
[0100] In this embodiment, through the annotation image with a learning annotation box in the training sample, the center point position, length, and width information of the learning annotation box can be accurately obtained, and the loss function corresponding to each annotation image can be accurately calculated through the center point position, length, and width information of the learning annotation box. Correspondingly, the loss function can be defined as:
[0101] Li
[0102] =P(Bxi)(Bxi - bxi) 2 +P(Byi)(Byi - byi)2 +P(Bli)(Bli - bli) 2 +P(Bwi)(Bwi - bwi) 2
[0103] Wherein, Li is the loss function corresponding to the i-th annotated image. P(Bxi), P(Byi), P(Bli), and P(Bwi) can be the probabilities corresponding to the generation of Bxi, Byi, Bli, and Bwi respectively, or the probability weighted values corresponding to the generation of Bxi, Byi, Bli, and Bwi respectively. Bxi is the abscissa of the center point of the i-th learning annotation box, Byi is the ordinate of the center point of the i-th learning annotation box, Bli is the length of the long side of the i-th learning annotation box, and Bwi is the length of the short side of the i-th learning annotation box.
[0104] Implementing the above embodiments has the following effects:
[0105] The technical solution of the present invention can achieve the situation where the result does not converge when the broken screen recognition model is training on the screen image annotated with the learning annotation box through the learning annotation box after data augmentation. At the same time, it can also effectively reduce the overfitting risk of the neural network model brought by the traditional rectangular annotation box with too large a range and inaccurate annotation.
[0106] Embodiment Three
[0107] Please refer to Figure 7 , which is a broken screen recognition method provided by the present invention and is implemented by the broken screen recognition model described in any one of the above Embodiment Two, including the following steps S301 - S303:
[0108] Step S301: Obtain the screen image to be recognized.
[0109] In this embodiment, the screen of the electronic terminal device to be recognized or detected can be photographed by a camera device to obtain the screen image to be recognized or detected of the electronic terminal device screen, or a real-time transmitted video stream. Wherein, the video stream is composed of one image frame after another. Therefore, when the obtained data is video stream data, the video stream data can also be extracted according to a preset frame interval to obtain at least one frame of image data, and the image frames that can recognize the screen in the extracted image frames can be recognized through the screen recognition model, so as to determine the final image frame as the screen image to be recognized according to the evaluation weights such as the integrity of the screen, the ability to recognize the crack area, and the image clarity, so as to accurately obtain the screen image to be recognized. Among them, the preset frame interval can be set according to the actual situation.
[0110] In this embodiment, the screen recognition model can be obtained by training a conventional neural network model based on sample data containing screen images of the terminal device, so as to realize the screen recognition of each image frame in the time-frequency stream and determine the screen images to be recognized that can be used for broken screen recognition.
[0111] Furthermore, the screen images to be recognized can include images of the complete terminal device or images that only contain part of the screen of the terminal device.
[0112] Step S302: Input the screen image to be recognized into the broken screen recognition model, so that the broken screen recognition model performs broken screen recognition on the screen image to be recognized; the broken screen recognition model is obtained by the construction method of the broken screen recognition model according to any one of Embodiment 2.
[0113] In this embodiment, the obtained screen image to be recognized is used as input data and input into a pre-trained broken screen recognition model. Among them, the broken screen recognition model is obtained by the construction method of the broken screen recognition model in Embodiment 2. After the screen image to be recognized is input, the broken screen recognition model will analyze and process the image. The broken screen recognition model will extract the features of the image and use the trained weight parameters for calculation to determine whether there is a broken screen area.
[0114] In this embodiment, the broken screen recognition model constructed in Embodiment 2 is a computer vision model constructed based on a deep learning algorithm, which can be used for tasks such as classifying screen images and recognizing and detecting broken screen targets, so that the broken screen recognition model can accurately and efficiently detect the broken areas on the screen.
[0115] Step S303: When a broken screen is recognized, generate a bounding box marked with the broken screen area on the screen image to be recognized.
[0116] In this embodiment, when the broken screen recognition model recognizes the broken screen area in the screen image to be recognized, a bounding box will be generated on the screen image to be recognized to mark the position and range of the broken screen area. Among them, since the training of the model and the training samples input into the model for learning are all marked by rectangular boxes, the bounding box in this embodiment can also be marked and represented by a rectangular box.
[0117] As another implementation, when no broken screen is recognized, there is no need to generate a bounding box on the screen image to be recognized, and the screen image to be recognized is added with a label of no broken screen.
[0118] As another embodiment, when generating a bounding box marked with the cracked screen area on the screen image to be recognized, a bounding box corresponding to the same size area as the learning bounding box in the training sample can be directly generated, and the screen image to be recognized with the marked bounding box can be output. As another embodiment, when the cracked screen area in the screen image to be recognized is recognized, all the cracked screen areas in the screen image to be recognized can be recognized, and a bounding box corresponding to the size of the first bounding box can be used to mark all the cracked screen areas, that is, by using the bounding box marking method that covers all the cracked screen areas to mark all the cracked screen areas, so as to realize the range covered by the complete cracked screen area, improve the integrity of the cracked screen recognition output result, avoid the problem that the marked cracked screen range is too limited, and improve the user experience. It can be understood that the learning bounding box of the training sample in the embodiments of the present invention can be used to improve the accuracy of model training and cracked screen recognition. However, due to data augmentation, the data marked by the bounding box is incomplete and inaccurate from the user's perspective. In order to enable the user to intuitively and accurately obtain the entire cracked screen recognition and marking result, the bounding box corresponding to the marking method of the present invention can be used for marking and recognition in the recognition stage. When the cracked screen recognition model of the present invention recognizes the cracked screen area in the screen image to be recognized, the bounding box range corresponding to the first bounding box is directly used for marking the cracked screen, and then the output of the cracked screen recognition result is realized.
[0119] It can be understood that by using the cracked screen recognition model, the cracked screen recognition of the screen image to be recognized can be automatically performed without manual intervention, which improves the recognition efficiency and accuracy. At the same time, when a cracked screen is recognized, a bounding box marked with the cracked screen area is generated, which can quickly and accurately locate the cracked position, making it very helpful for subsequent operations such as repair or screen replacement. Moreover, due to the use of the trained cracked screen recognition model, the possibility of misjudgment can be reduced. Compared with manual judgment, the model can more objectively analyze the image features and improve the recognition accuracy.
[0120] Implementing the above embodiments has the following effects:
[0121] Compared with the traditional manual detection method, the technical solution of the embodiments of the present invention can greatly improve the detection efficiency. It can not only process multiple screen images to be recognized simultaneously, but also complete the recognition of the entire screen in a short time. At the same time, the generated bounding box can be used to record and analyze the cracked screen situation. Through the statistics and analysis of a large amount of data, some valuable conclusions can be obtained, such as the frequency of cracked screens, location distribution and other information, which provides a reference basis for subsequent repair and improvement and improves the user experience.
[0122] Embodiment 4
[0123] Please refer to Figure 8, which is a labeling device provided by the present invention, including: a screen image module 110, a sampling module 120, and a labeling module 130.
[0124] The screen image module 110 is configured to obtain a screen image labeled with a first labeling frame and a second labeling frame; wherein, the first labeling frame includes a cracked screen extension area of the screen, and the second labeling frame includes a cracked screen core area of the screen;
[0125] The sampling module 120 is configured to sample a probability distribution of an initial learning labeling frame according to a distance relationship between the borders of the first labeling frame and the second labeling frame.
[0126] The labeling module 130 is configured to calculate a final learning labeling frame according to the probability distribution of the initial learning labeling frame and label it on the screen image.
[0127] As a preferred solution, the sampling of the probability distribution of the initial learning labeling frame according to the distance relationship between the borders of the first labeling frame and the second labeling frame includes:
[0128] Obtain the number of rows and columns of pixel values occupied by a horizontal side and a vertical side in the second labeling frame in the screen image respectively, and take the number of rows as the first mean value and the number of columns as the second mean value; calculate the distances that all sides of the second labeling frame are respectively translated and extended to the corresponding sides of the first labeling frame according to the positions of the first labeling frame and the second labeling frame in the screen image; wherein, the screen image area covered by the second labeling frame is within the screen image area covered by the first labeling frame, and each side of the second labeling frame corresponds to one side of the first labeling frame; scale the distances of the extended sides of all sides according to a preset ratio as the variance corresponding to the side; wherein, each side corresponds to a variance; generate a probability distribution of each side of the initial learning labeling frame according to the first mean value, the second mean value, and the variance corresponding to each side.
[0129] As a preferred solution, the initial learning labeling frame includes a first side, a second side, a third side, and a fourth side; the first side and the second side are respectively the horizontal side and the vertical side forming an angle in the initial learning labeling frame, and the third side and the fourth side are respectively the horizontal side and the vertical side forming the other diagonal in the initial learning labeling frame;
[0130] The length of the first side is greater than the first mean value, the length of the second side is greater than the second mean value, the length of the third side is less than the first mean value, and the length of the fourth side is less than the second mean value.
[0131] As a preferred solution, calculating a final learned annotation box according to the probability distribution of the initial learned annotation box and annotating it on the screen image includes:
[0132] Calculating the probability of the center point position, the length probability, and the width probability of the final learned annotation box according to the probability distribution of each side of the initial learned annotation box;
[0133] Obtaining the final learned annotation box according to the center point position probability, the length probability, and the width probability, and annotating the final learned annotation box on the screen image.
[0134] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the above-described device can refer to the corresponding process in the foregoing method embodiment, and will not be elaborated herein.
[0135] Implementing the above embodiments has the following effects:
[0136] The technical solution of the present invention enables quick positioning to the cracked screen extension area and the cracked screen core area through a screen image marked with the distance relationship between the first annotation box and the second annotation box, and samples to obtain the probability distribution of the initial learned annotation box, and then calculates the final learned annotation box and annotates it on the screen image. The annotation of the cracked screen extension area and the cracked screen core area can sample and calculate the true value of the learned annotation box, thereby realizing data enhancement of the annotation box, making the sampling result slightly perturbed within a certain range to improve the accuracy of the annotation data, and avoiding the problem of large errors in the learned annotation box used for model training.
[0137] Embodiment Five
[0138] Please refer to Figure 9 , which is a device for constructing a cracked screen recognition model provided by the present invention, including: a sample module 210 and a training module 220.
[0139] The sample module 210 is used to obtain training samples; wherein, the training samples include at least one original screen image and at least one annotation image marked with a learned annotation box corresponding to the original screen image.
[0140] The training module 220 is used to construct a preset neural network and input the training samples into the neural network for training to obtain a cracked screen recognition model.
[0141] Among them, the annotation image of the training sample in the method for constructing the cracked screen recognition model is obtained by the annotation method described in the above Embodiment One.
[0142] As a preferred solution, the steps of constructing a preset neural network and inputting the training samples into the neural network for training to obtain a broken screen recognition model are as follows:
[0143] Construct a preset neural network;
[0144] Based on the labeled images in the training samples with learning labeled frames, obtain the center point position, length, and width information of the learning labeled frames;
[0145] According to the center point position, length, and width information of the learning labeled frames in each labeled image, calculate the loss function corresponding to each labeled image;
[0146] Input the training samples into the neural network and combine the loss function corresponding to each labeled image to train and obtain a broken screen recognition model.
[0147] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the above-described device can refer to the corresponding process in the foregoing method embodiment and will not be elaborated herein.
[0148] Implementing the above embodiments has the following effects:
[0149] The technical solution of the present invention can solve the problem that the result of the broken screen recognition model does not converge when training on screen images labeled with learning labeled frames through the data-augmented learning labeled frames, and can also effectively reduce the overfitting risk of the neural network model caused by the traditional rectangular labeled frames with too large a range and inaccurate labeling.
[0150] Embodiment Six
[0151] Please refer to Figure 10 , which is a broken screen recognition device provided by the present invention, including: a screen image module 310, a recognition module 320, and an output module 330;
[0152] The screen image module is used to obtain a screen image to be recognized;
[0153] The recognition module is used to input the screen image to be recognized into the broken screen recognition model so that the broken screen recognition model performs broken screen recognition on the screen image to be recognized;
[0154] The output module is used to generate a labeled frame labeled with a broken screen area on the screen image to be recognized when a broken screen is recognized;
[0155] Wherein, the broken screen recognition model is constructed by the construction method of the broken screen recognition model in any one of the above items, and the labeled images of the training samples in the construction method of the broken screen recognition model are obtained by the labeling method in any one of the above items.
[0156] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working process of the above-described device can refer to the corresponding process in the foregoing method embodiments and will not be elaborated herein.
[0157] Implementing the above embodiments has the following effects:
[0158] Compared with the traditional manual detection method, the technical solution of the embodiments of the present invention can greatly improve the detection efficiency. It can not only process multiple images of the screen to be recognized simultaneously, but also complete the recognition of the entire screen in a short time. At the same time, the generated annotation boxes can be used to record and analyze the broken screen situation. Through the statistics and analysis of a large amount of data, some valuable conclusions can be obtained, such as the frequency and location distribution of broken screens, etc., providing a reference basis for subsequent maintenance and improvement, and improving the user experience.
[0159] Embodiment Seven
[0160] Correspondingly, the present invention further provides a terminal device, including: a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the annotation method described in any one of the above Embodiment One, the method for constructing a broken screen recognition model described in any one of the above Embodiment Two, or the broken screen recognition method described in the above Embodiment Three.
[0161] The terminal device of this embodiment includes: a processor, a memory, and a computer program and computer instructions stored in the memory and executable on the processor. When the processor executes the computer program, it implements each step in the above Embodiment One, such as Figure 3 the steps S101 to S103 shown, or Figure 6 the steps S201 - S202 shown, or Figure 7 the steps S301 - S303 shown. Alternatively, when the processor executes the computer program, it implements the functions of each module / unit in the above device embodiment, such as the screen image module 110, the sampling module 120, and the annotation module 130, or the sample module 210 and the training module 220, or the screen image module 310, the recognition module 320, and the output module 330.
[0162] Exemplarily, the computer program can be divided into one or more modules / units. The one or more modules / units are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the terminal device.
[0163] The terminal device may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The terminal device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the schematic diagram is only an example of the terminal device and does not constitute a limitation on the terminal device. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the terminal device may also include input / output devices, network access devices, a bus, etc.
[0164] The so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the terminal device, connecting various parts of the entire terminal device through various interfaces and lines.
[0165] The memory may be used to store the computer program and / or module. The processor realizes various functions of the terminal device by running or executing the computer program and / or module stored in the memory, and by calling the data stored in the memory. The memory may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function, etc.; the data storage area may store data created according to the use of the mobile terminal, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0166] Among them, if the modules / units integrated in the terminal device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present invention, it can also be completed by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0167] Embodiment Eight
[0168] Correspondingly, the present invention further provides a computer-readable storage medium. The computer-readable storage medium includes a stored computer program. Among them, when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the annotation method described in any one of the above Embodiment One, the method for constructing a cracked screen recognition model described in any one of the above Embodiment Two, or the cracked screen recognition method described in the above Embodiment Three.
[0169] The above-mentioned specific embodiments have further elaborated on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above-mentioned are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. It is particularly pointed out that for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A labeling method, characterized in that, Including: Obtain a screen image marked with a first annotation box and a second annotation box; wherein, the first annotation box includes a cracked screen extension area of the screen, and the second annotation box includes a cracked screen core area of the screen; Sample a probability distribution of an initial learning annotation box according to a distance relationship between the borders of the first annotation box and the second annotation box; Calculate a final learning annotation box according to the probability distribution of the initial learning annotation box, and mark it on the screen image.
2. The marking method according to claim 1, characterized in that, The sampling of the probability distribution of the initial learning annotation box according to the distance relationship between the borders of the first annotation box and the second annotation box includes: Obtain the number of rows and columns of pixel values occupied by a horizontal side and a vertical side in the second annotation box in the screen image respectively, and use the number of rows as the first mean value and the number of columns as the second mean value; According to the positions of the first annotation box and the second annotation box in the screen image, calculate the distances that all sides of the second annotation box are respectively translated and extended to the corresponding sides of the first annotation box; wherein, the screen image area covered by the second annotation box is within the screen image area covered by the first annotation box, and each side of the second annotation box corresponds to one side of the first annotation box; Scale the extended distances of all sides according to a preset ratio as the variance corresponding to each side; wherein, each side corresponds to one variance; Generate a probability distribution of each side of the initial learning annotation box according to the first mean value, the second mean value and the variance corresponding to each side.
3. The marking method according to claim 2, characterized in that, The initial learning annotation box includes a first side, a second side, a third side and a fourth side; the first side and the second side are respectively the horizontal side and the vertical side forming an angle in the initial learning annotation box, and the third side and the fourth side are respectively the horizontal side and the vertical side forming the other diagonal angle in the initial learning annotation box; The length of the first side is greater than the first mean value, the length of the second side is greater than the second mean value, the length of the third side is less than the first mean value, and the length of the fourth side is less than the second mean value.
4. A marking method according to any one of claims 1-3, characterized in that, The calculating of the final learning annotation box according to the probability distribution of the initial learning annotation box and marking it on the screen image includes: Calculate the center point position probability, length probability and width probability of the final learning annotation box according to the probability distribution of each side of the initial learning annotation box; Obtain the final learning annotation box according to the center point position probability, length probability and width probability, and mark the final learning annotation box on the screen image.
5. A method for constructing a broken screen recognition model, characterized in that, The method includes: Obtain training samples; wherein, the training samples include at least one original screen image and at least one annotated image marked with a learning annotation box corresponding to the original screen image; the annotated image marked with the learning annotation box is obtained by the annotation method according to any one of claims 1-4; Construct a preset neural network, and input the training samples into the neural network for training to obtain a cracked screen recognition model.
6. The method for constructing a broken screen recognition model according to claim 5, wherein, Construct a preset neural network, and input the training samples into the neural network for training to obtain a broken screen recognition model, specifically as follows: Construct a preset neural network; According to the labeled images with learning labeled frames in the training samples, obtain the center point position, length, and width information of the learning labeled frames; According to the center point position, length, and width information of the learning labeled frames of each labeled image, calculate the loss function corresponding to each labeled image; Input the training samples into the neural network, and combine with the loss function corresponding to each labeled image to train and obtain a broken screen recognition model.
7. A broken screen recognition method, characterized in that, The method includes: Obtain an image of the screen to be recognized; Input the image of the screen to be recognized into the broken screen recognition model, so that the broken screen recognition model performs broken screen recognition on the image of the screen to be recognized; the broken screen recognition model is obtained by the method for constructing the broken screen recognition model according to any one of claims 5-6; When a broken screen is recognized, generate a labeled frame labeled with the broken screen area on the image of the screen to be recognized.
8. A labeling device, characterized in that, It includes: A screen image module, a sampling module, and a labeling module; The screen image module is used to obtain a screen image labeled with a first labeled frame and a second labeled frame; wherein, the first labeled frame includes the broken screen extension area of the screen, and the second labeled frame includes the broken screen core area of the screen; The sampling module is used to sample the probability distribution of the initial learning labeled frame according to the distance relationship between the borders of the first labeled frame and the second labeled frame; The labeling module is used to calculate the final learning labeled frame according to the probability distribution of the initial learning labeled frame and label it on the screen image.
9. An apparatus for constructing a broken screen recognition model, characterized in that It includes: A sample module and a training module; The sample module is used to obtain training samples; wherein, the training samples include at least one original screen image and at least one labeled image labeled with a learning labeled frame corresponding to the original screen image; The training module is used to construct a preset neural network, and input the training samples into the neural network for training to obtain a broken screen recognition model; Wherein, the labeled image of the training sample in the method for constructing the broken screen recognition model is obtained by the labeling method according to any one of claims 1-4.
10. A broken screen recognition device, characterized in that, It includes: A screen image module, a recognition module, and an output module; The screen image module is used to obtain an image of the screen to be recognized; The recognition module is used to input the image of the screen to be recognized into the broken screen recognition model, so that the broken screen recognition model performs broken screen recognition on the image of the screen to be recognized; The output module is used to generate a labeled frame labeled with the broken screen area on the image of the screen to be recognized when a broken screen is recognized; Wherein, the broken screen recognition model is constructed by the method for constructing the broken screen recognition model according to any one of claims 5-6, and the labeled image of the training sample in the method for constructing the broken screen recognition model is obtained by the labeling method according to any one of claims 1-4.
11. A terminal device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the annotation method according to any one of claims 1 to 4, the method for constructing a broken screen recognition model according to any one of claims 5 to 6, or the broken screen recognition method according to claim 7.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program. When the computer program runs, it controls the device where the computer-readable storage medium is located to execute the annotation method according to any one of claims 1 to 4, the method for constructing a broken screen recognition model according to any one of claims 5 to 6, or the broken screen recognition method according to claim 7.