Page layout and element analysis method based on deep learning

By improving the YOLO and PaddleOCR models and combining them with K-means++ and DBSCAN algorithms, the accuracy and fusion problems of set-top box page element recognition were solved, achieving efficient and accurate page layout analysis, improving recognition accuracy and information fusion capabilities, and supporting page interaction and design optimization.

CN120913191APending Publication Date: 2025-11-07BEIJING BOHUI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511006492.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Due to differences in the layout of set-top box pages among different operators, the accuracy of object detection and character recognition is poor, and the detection and recognition results cannot be effectively integrated. There is a lack of multi-level analysis of elements, making it difficult to meet the needs of efficient and accurate page element analysis.

Method used

We employ a deep learning-based approach, improving the YOLO model for object detection and identifying image element regions. We also combine the PaddleOCR model for character recognition, merging text and image bounding boxes. Finally, we use K-means++ and DBSCAN algorithms for clustering and grouping to analyze the layout patterns of the elements.

Benefits of technology

It improves the accuracy and integration of set-top box page element recognition, enables efficient understanding and organization of page layout, enhances recognition accuracy and information fusion capabilities, and supports page interaction and design optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913191A_ABST
    Figure CN120913191A_ABST
Patent Text Reader

Abstract

The invention discloses a page layout and element analysis method based on deep learning, and aims to solve the problem that an existing target detection and character recognition method generally cannot meet the requirement for efficient and accurate analysis of page elements due to the complexity of a set top box page. According to the application, target detection is carried out on a set top box page image based on a first deep learning model, an image element region in a page is identified, a plurality of corresponding image frames are generated, and a live broadcast window region and a poster recommendation position region are determined according to the area of the image frames, the rest image frames are poster recommendation position areas; recognizing characters in the set top box page image based on a second deep learning model, and generating a plurality of textboxes; and merging the overlapped or adjacent original elements by judging the overlapping condition and the distance relationship between the textboxes and the overlapping condition of the textboxes and the image box.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of target detection, character recognition and image processing, and particularly relates to a page layout and element analysis method. BACKGROUND

[0002] With the development of digital television technology, set-top boxes have become an important part of people's daily life. However, some broadcast television operators have irregular behaviors, such as non-transparent charging structure, users need to pay additional fees for different channels and services; cumbersome operation, non-intuitive navigation, especially poor experience for the elderly and users unfamiliar with technology. In order to regulate the content provided by the operator, the State Administration of Radio, Film and Television issued the "Cable Television, Interactive Network Television (IPTV) and Internet Television Page, Charging and Operation Management Specification (Trial)", which proposes to solve the problems of "nested doll" charging and complex operation. However, the differences in technical architecture and interface design of different operators increase the difficulty of element detection and recognition, and the implementation of multi-dimensional automatic checking still faces great challenges.

[0003] The output interface of the set-top box is usually complex in layout, composed of multiple interactive components, such as live window area, status bar, navigation bar, poster recommendation position, etc., each component is presented in the form of characters and images. Image components can locate object regions through target detection technology, and character components can identify the text content therein through character recognition technology. At present, target detection and character recognition technology has been widely used in image and text processing fields, but the complexity of the set-top box page makes the existing target detection and character recognition methods usually unable to meet the needs of efficient and accurate analysis of page elements.

[0004] The inventors have realized that the prior art has the following problems in particular:

[0005] 1) Poor accuracy of target detection and character recognition: the set-top box page has great differences in layout due to different operators, resulting in different scales of live window area and poster recommendation position, and some characters are presented in the form of inclined and artistic fonts, which will increase the difficulty of target detection and character recognition;

[0006] 2) Target detection and character recognition results cannot be effectively fused: character recognition can recognize the text in the page, but cannot directly match the live window area, poster recommendation position area and other regions of target detection, resulting in poor relevance of text information and object information, which will increase the difficulty of later page layout analysis;

[0007] 3) Lack of multi-level analysis of elements: although the set-top box page is complex and diverse in layout, there are still certain arrangement rules, and the existing technology often only focuses on single tasks of target detection or character recognition, lacking hierarchical analysis of the results. SUMMARY

[0008] The application provides a page layout and element analysis method based on deep learning, aiming to solve the problem that the complexity of the set-top box page makes the existing target detection and character recognition method generally unable to meet the demand for efficient and accurate analysis of page elements.

[0009] In a first aspect, a page layout and element analysis method based on deep learning is provided, comprising:

[0010] Target detection: based on a first deep learning model, the set-top box page image is detected to identify the image element area in the page, and a plurality of image frames are generated. According to the area size of the image frame, the live window area and the poster recommendation position area are determined: the image frame with the largest area is the live window area, and the remaining image frames are the poster recommendation position area;

[0011] Character recognition: based on a second deep learning model, the characters in the set-top box page image are recognized to generate a plurality of text frames; the first deep learning model and the second deep learning model are constructed in the training stage by pulling the code stream from the set-top box page of different operators to obtain the real set-top box page image to construct the data set;

[0012] Element merging: each image frame and text frame is regarded as a separate original element. By judging the overlapping condition, distance relationship between the text frames and the overlapping condition between the text frames and the image frames, the overlapping or adjacent original elements are merged, wherein one original element is regarded as the main element and the other original elements are regarded as the subsidiary elements. For the merged elements, the subsidiary attributes of the main elements are determined according to the contents of the subsidiary elements;

[0013] Clustering grouping: according to the size and position characteristics of the merged elements, clustering processing is performed, and according to the clustering result, the elements are grouped, including the status bar group, the navigation bar group and the poster recommendation position group. The alignment mode of each element group is recorded, the relative position relationship of the element groups is analyzed, and the layout mode of the element page is determined.

[0014] In the above scheme, further optionally, the first deep learning model is an improved YOLO model, specifically using a K-means++ clustering algorithm to dynamically select suitable anchor points according to data distribution.

[0015] In the above scheme, further optionally, the training process of the first deep learning model comprises:

[0016] S11, using an image labeling tool LabelImg to label a live window region and a poster recommendation position region of each image in a set of image data of a set-top box page, using a minimum circumscribed rectangle method to label a boundary region of each target, and a label being rectangle;

[0017] S12, obtaining real boundary sizes of all labeled boxes from a label file derived after labeling, to obtain width and height values of all labeled boxes;

[0018] S13, randomly dividing image data and label data into a first training set, a first verification set and a first test set according to a first preset ratio;

[0019] S14, using K-means++ to cluster the labeled boxes to obtain new anchor points;

[0020] S15, replacing the generated anchor point data into a configuration file of a YOLO model to obtain an improved YOLO model, training the improved YOLO model using the first training set, verifying model performance using the first verification set during the training process, evaluating model performance using the first test set after model training and verification are completed, and obtaining the first deep learning model.

[0021] In the above scheme, further optionally, the step S14 specifically includes:

[0022] S141, selecting an initial clustering center: randomly selecting a data point from all labeled boxes as a first clustering center;

[0023] S142, calculating Euclidean distance: for each remaining data point, calculating a Euclidean distance between the data point and a nearest selected clustering center, and a specific formula being:

[0024]

[0025] wherein d(x i ,c i ) represents the Euclidean distance, k represents a dimension index, n represents a dimension of a space, x i represents a data point, c i represents a clustering center; x ik represents the data point in the k dimension; c ik represents the clustering center in the k dimension;

[0026] S143, updating the clustering center: according to a distance of each data point to a nearest clustering center, selecting a next clustering center in a probabilistic manner, and a probability of selecting the data point x i as the next clustering center being proportional to a square of the distance, and a specific formula being:

[0027]

[0028] wherein P(x i ) represents a probability of selecting data point x i , d(x i , c i ) represents a distance of data point x i to the nearest cluster center c i , X represents a set of all data points, x j represents the jth data point, and d(x j, , c i ) represents a distance of the jth data point to the nearest cluster center c i ;

[0029] S144, repeating step S142 and step S143 until a plurality of cluster centers are selected;

[0030] S145, performing a standard K-means clustering, assigning each data point to the nearest cluster center, and then updating the cluster center until the clustering result converges, obtaining the final plurality of cluster centers as anchor points; wherein the number of cluster centers is determined according to the needs of the model.

[0031] In the above scheme, optionally, the second deep learning model enhances the recognition ability of special fonts by optimizing hyperparameters.

[0032] In the above scheme, further optionally, the training process of the second deep learning model specifically includes:

[0033] S21, using PaddleOCR labeling tool to label the boundary region of characters in the set-top box page image data, including:

[0034] S211, using the pre-training model of PaddleOCR and the labeling tool PPOCRLabel to automatically label the character region of the set-top box page image data, and initially generating a labeling result;

[0035] S212, checking the labeling situation one by one, manually modifying the mislabeled and missed labeled characters, and ensuring the accuracy of the labeling result;

[0036] S213, specifying the specific rotation angle for each text region to ensure that it can adapt to text with different inclination angles and improve the adaptability of the model to complex text;

[0037] S22, after completing the labeling of the character region, exporting the labeling result, and dividing the labeled image and the corresponding label into a second training set, a second validation set and a second test set according to a second preset proportion;

[0038] S23, download the detection model, recognition model and direction classification model of the PaddleOCR model and the respective configuration files, and perform model fine-tuning on this basis;

[0039] S24, model fine-tuning: training the detection model, recognition model and direction classification model of the PaddleOCR model using the second training set, verifying the performance of each model using the second verification set, and adjusting the hyperparameters corresponding to each model according to the verification results;

[0040] S25, after the model training and verification are completed, the performance of the model is evaluated using the second test set, and the second deep learning model is obtained.

[0041] In the above scheme, the element merging step includes:

[0042] S31, traverse all detected elements, and store the text box and the image box separately; wherein the text box is obtained through the character recognition step, and the image box is obtained through the target detection step;

[0043] S32, text box merging, including:

[0044] S321, sort the text boxes according to the horizontal coordinate of the upper left corner;

[0045] S322, traverse each text box and check the text boxes adjacent or overlapping with it;

[0046] S323, calculate the intersection area of the text box being checked and other text boxes adjacent or overlapping with it;

[0047] S324, if the intersection area of the two text boxes is greater than 0, merge them into one text box, if there are multiple text boxes with intersection area greater than 0, only merge the two text boxes with the largest intersection area, and the merged text box no longer participates in the subsequent merging process; wherein the intersection area of multiple text boxes greater than 0 means that there are more than two text boxes overlapping (without adjacent), at this time only the two text boxes with the largest overlapping area are merged.

[0048] S325, when merging, update the position and text content of the merged text box; the updated position is the smallest closed region containing both text boxes, and the text content is the text content of the text box with larger area, and the removed text box content is recorded as an attached attribute of the preserved text content;

[0049] S33, text box distance merging, including:

[0050] S331, for the remaining unmerged text boxes, check the distance of the adjacent text boxes;

[0051] S332, calculate the horizontal distance and vertical distance between the current text box and the adjacent text box, if the horizontal distance or the vertical distance between the two text boxes is less than a preset number of pixels, then merge them into one text box; if the horizontal distance or the vertical distance between more than two text boxes is less than a preset number of pixels, then only merge the two text boxes with the smallest distance; the merged text box no longer participates in the subsequent merging process; wherein, when judging whether two text boxes are adjacent, the preset pixel distance is used, if the current text box has multiple adjacent text boxes with a distance less than the preset pixel, then select the text box with the smallest horizontal distance or vertical distance from the current text box, and merge the text box with the current text box;

[0052] S333, when merging, update the position and text content of the merged text box; the updated position is the smallest closed region containing both text boxes, and the text content is the text content of the text box with larger area, and the removed text box content is recorded as an attached attribute of the reserved text content;

[0053] S34, text box and image box merging, including:

[0054] S341, for the remaining text boxes that do not participate in merging, check whether they have intersection with the image box;

[0055] S342, if the intersection area of the text box and the image box is greater than 0, then remove such text box, and record the removed text box content as an attached attribute of the image box.

[0056] In the above scheme, optionally, the attached attribute of the main element includes a charging attribute; the content of the attached element includes a keyword about whether the user needs to pay, and the charging attribute of the main element is determined according to the keyword.

[0057] In the above scheme, optionally, before clustering and grouping, element optimization is further performed, including:

[0058] S41, remove text boxes with text length less than or equal to 1 and meaningless identifier;

[0059] S42, for the merged elements, reassign element ID to ensure that the element ID starts from 1 and gradually increases, avoiding empty ID or repeated ID.

[0060] In the above scheme, further optionally, the step of clustering and grouping specifically includes:

[0061] S51, according to the size and position characteristics of the merged elements, clustering is performed using the DBSCAN algorithm to identify groups of elements with similar sizes, specifically including:

[0062] S511, an unvisited element is randomly selected from the elements as the current processing point P of the DBSCAN algorithm, and an optimized element is regarded as a point;

[0063] S512, the ε neighborhood of the point P is checked, and the number of unvisited points contained in the neighborhood except for the point P is calculated, if the number of points in the neighborhood is greater than or equal to the minimum boundary point number of the ε neighborhood, the point P is regarded as a core point; wherein the minimum boundary point number is the parameter min_samples in the DBSCAN algorithm;

[0064] S513, all points in the neighborhood of the core point are added to the cluster, and for each point in the neighborhood, if it is also a core point, the neighborhood of the point is continuously expanded, and the recursion is continued until no new core point can be expanded;

[0065] S514, noise processing: if a point is not a core point and does not belong to the neighborhood of any core point, it is marked as noise and removed;

[0066] S515, steps S511-S514 are performed on all points until all points are visited;

[0067] S52, the preliminary clustering result is corrected based on the size and position characteristics of the elements, and it is checked whether adjacent elements belong to the same group, if there is obvious connection across groups, the groups are merged to ensure that the distance or size change between elements in the same group is not too large, if significant inconsistency is found, the clustering division is adjusted;

[0068] S53, elements are grouped according to the clustering result, and the alignment mode of each element group is further recorded; the relative position relationship of the element groups is analyzed to identify which elements are horizontally arranged and which elements are vertically arranged to ensure the accuracy of the layout relationship;

[0069] S54, elements causing clustering conflicts are re-assigned, and the assignment method includes:

[0070] S541, average area judgment: the average area of the element is calculated, and the element is assigned to the group with larger area to ensure that each element is reasonably classified;

[0071] S542, relative position adjustment: if the element is located at the junction of multiple clusters, the final attribution of the element is determined by the relative position and size attribute of the element;

[0072] S55, by comparing the layout features of multiple element groups, a layout pattern with similar structure is identified;

[0073] S56, for isolated elements not participating in the initial clustering, analyze the spatial gap between unclustered elements and clustered elements, if the distance is close or has obvious relative position relationship, the unclustered element is re assigned to the corresponding group.

[0074] Compared with the prior art, the present application has at least the following beneficial effects:

[0075] Based on further analysis and research on the problems of the prior art, it is realized that the complexity of the set-top box page makes the existing target detection and character recognition method generally unable to meet the demand for efficient and accurate analysis of page elements, which is specifically manifested in: (1) poor accuracy of target detection and character recognition; (2) the results of target detection and character recognition cannot be effectively combined; (3) lack of multi-level analysis of elements. Through deep learning technology, the present application optimizes the hyperparameters of the detection and recognition model, enhances the recognition ability of artistic fonts or inclined text, and improves the accuracy and robustness of character recognition; the detected text box and image box are merged and processed, the element integration is optimized, and repeated recognition is avoided, which improves the accuracy and integration of page element recognition; the clustering algorithm is used to cluster the elements, accurately analyze the page layout, and realize efficient understanding and organization of the set-top box page structure. Thus, the accuracy, integration and efficiency of set-top box page layout and element analysis are improved, and a better data basis and analysis result for subsequent processing and application of the set-top box page are provided.

[0076] The present application combines traditional vision and deep learning technology to propose an efficient and accurate set-top box page element detection and layout analysis method. This method can significantly improve the recognition accuracy and layout analysis ability of set-top box page elements, and the specific effects are as follows:

[0077] 1. Improve recognition accuracy: through the improved YOLO model and fine-tuned PaddleOCR model, the present application can effectively recognize various elements in the set-top box page, especially for complex text (such as artistic fonts, inclined fonts, etc.) and diversified layout, the recognition accuracy is greatly improved.

[0078] 2. Effective information fusion: this method combines target detection and character recognition technology, not only can recognize objects in the set-top box page (such as live window, poster recommendation position, etc.), but also can accurately extract text information and effectively fuse it with image elements, solving the problem that the results of target detection and character recognition in the prior art cannot be effectively combined.

[0079] 3. Hierarchical layout analysis: Through multi-level analysis of elements, the invention can provide more comprehensive information for page layout, helping to better understand the structure and design rules of set-top box pages. This not only helps page interaction and design optimization, but also provides support for subsequent user experience analysis.

[0080] 4. Wide applicability: This technology is not only suitable for set-top box page element detection and layout analysis, but also can be extended to other complex user interface recognition and optimization tasks, with high universality and application prospects. BRIEF DESCRIPTION OF DRAWINGS

[0081] Figure 1 The flowchart of the page layout and element analysis method based on deep learning provided by the first embodiment of the present application is shown;

[0082] Figure 2 The flowchart of the page layout and element analysis method based on deep learning provided by the second embodiment of the present application is shown;

[0083] Figure 3 The improved YOLO target detection effect diagram provided by an embodiment of the present application is shown;

[0084] Figure 4 The PaddleOCR character recognition effect diagram provided by an embodiment of the present application is shown;

[0085] Figure 5 The element merging and optimization effect diagram provided by an embodiment of the present application is shown;

[0086] Figure 6 The clustering grouping diagram provided by an embodiment of the present application is shown. DETAILED DESCRIPTION

[0087] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application will be further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.

[0088] In the description of the present application: unless otherwise specified, the meaning of "multiple" is two or more. The terms "first", "second", "third" and the like in the present application are intended to distinguish the objects referred to, and do not have special technical connotations (for example, should not be understood as emphasizing importance or order, etc.). The expressions "include", "contain", "have" and the like also mean "not limited to" (some units, components, materials, steps, etc.).

[0089] The application provides a set-top box page element detection and layout analysis method based on traditional vision and deep learning, which solves the problems of low precision, poor information fusion and insufficient hierarchical analysis in the existing set-top box page element recognition. By combining target detection, character recognition and image processing technology, various elements in the set-top box page can be efficiently and accurately extracted, and hierarchical layout analysis can be performed, thereby improving the ability of page interaction and design optimization. The related technology can be applied to user interface recognition scenarios other than set-top box pages.

[0090] In one embodiment, referring to Figure 1 , a deep learning-based page layout and element analysis method is provided, including target detection, character recognition, element merging and clustering grouping.

[0091] Target detection: based on a first deep learning model, target detection is performed on the set-top box page image to identify image element regions in the page and generate a plurality of image frames, and according to the area size of the image frames, a live window region and a poster recommendation position region are determined: the image frame with the largest area is the live window region, and the remaining image frames are the poster recommendation position region;

[0092] Character recognition: based on a second deep learning model, characters in the set-top box page image are recognized to generate a plurality of text frames; the first deep learning model and the second deep learning model are both constructed by pulling code streams from set-top box pages of different operators to obtain real set-top box page images and build a data set during the training stage;

[0093] Element merging: each image frame and text frame is treated as a separate original element, and by judging the overlapping condition, distance relationship between text frames, and overlapping condition between text frames and image frames, overlapping or adjacent original elements are merged, wherein one original element is the main element and the other original elements are the subsidiary elements; for the merged elements, the subsidiary attributes of the main element are determined according to the contents of the subsidiary elements;

[0094] Clustering grouping: according to the size and position characteristics of the merged elements, clustering processing is performed, and according to the clustering results, the elements are grouped, including a status bar group, a navigation bar group, and a poster recommendation position group; the alignment mode of each element group is recorded, the relative position relationship of the element groups is analyzed, and the layout mode of the element page is determined.

[0095] To achieve high-precision element recognition and layout analysis of set-top box pages, the application realizes automatic recognition and hierarchical analysis of set-top box pages based on target detection, character recognition technology in deep learning and image processing methods in traditional vision, which not only accurately extracts page elements, but also performs hierarchical analysis and semantic understanding on these elements, and provides effective support for subsequent page interaction, design optimization and user experience analysis.

[0096] The application can efficiently and accurately identify various elements in the set-top box page and perform hierarchical analysis and semantic understanding by fusing target detection and character recognition technology in deep learning and image processing methods in traditional vision. Not only does it significantly improve the extraction and layout analysis accuracy of set-top box page elements, but it also provides strong technical support for the design and optimization of set-top boxes and other complex user interfaces. The flowchart of the entire method is shown in Figure 2

[0097] The method is aimed at the characteristics of set-top box page images. Starting from the input image, the improved YOLO model and the fine-tuned PaddleOCR model are used for target detection and character recognition, then the detection and recognition results are merged, and then the merged elements are optimized to improve the processing effect, and finally the optimized elements are clustered and grouped to complete the entire image processing and analysis process.

[0098] I. Improved target detection of YOLO

[0099] Based on the improved YOLO model, the live window and poster recommendation area in the set-top box page are identified, and the specific steps are as follows:

[0100] 1. Pull the code stream from the set-top box pages of different operators to obtain real set-top box page image data.

[0101] 2. Use LabelImg to label the live window area and poster recommendation area of each image in the data set. Use the minimum bounding rectangle method to label the boundary area of each target, and the label is "rectangle".

[0102] 3. Get the real boundary size of all labeled boxes from the labeled data set to get the width and height values of all labeled boxes.

[0103] 4. Use K-means++ to cluster the labeled boxes to get new anchor points:

[0104] (1) Select the initial clustering center. Randomly select a data point from the labeled data set as the first clustering center.

[0105] (2) Calculate the Euclidean distance. For each data point in the data set, calculate the Euclidean distance between it and the nearest selected clustering center, the specific formula is:

[0106]

[0107] ​wherein d(x i ,c i ) represents the Euclidean distance, k represents the dimension index, n represents the dimension of the space, x i represents the data point, c i represents the cluster center; x ik represents the data point in k dimensions; and c ik represents the cluster center in k dimensions.

[0108] (3) Update the cluster center. According to the distance of each data point to the nearest cluster center, the next cluster center is selected in a probabilistic manner, and the probability of selecting the data point x i as the next cluster center is proportional to the square of its distance, and the specific formula is:

[0109]

[0110] wherein P(x i ) represents the probability of selecting the data point x i , d(x i ,c i ) represents the distance of the data point x i to the nearest cluster center c i , X represents the set of all data points, x j represents the jth data point, and d(x j, c i ) represents the distance of the jth data point to the nearest cluster center c i .

[0111] Repeat steps (2) and (3) until m cluster centers are selected. Since the YOLO model requires 9 anchor points, the value of m is 9.

[0112] (4) Perform standard K-means clustering to assign each data point to the nearest cluster center, and then update the cluster center until the clustering result converges. Finally, the 9 cluster centers found, i.e., the anchor points, are obtained.

[0113] 5. Randomly divide the image data and label data into a first training set, a first validation set, and a first test set in a ratio of 8:1:1.

[0114] 6. Replace the generated anchor point data in the Anchors in the YOLOv5s.yaml configuration file, and then use the yolov5s model to train the data set to obtain a set-top box page detector. (The first training set is used to train the improved YOLO model, and the first validation set is used to verify the model performance during the training process. After the model training and verification are completed, the first test set is used to evaluate the performance of the model, and the first deep learning model is obtained.)

[0115] 7. Use the trained detector (first deep learning model) to detect images in the test set, and get the detection results, where the largest detection box is adjusted as the live window area, and the rest are poster recommendation area.

[0116] In this embodiment, the correspondence between technical terms and detection targets is as follows:

[0117] Data points: the width and height value pair of the annotation box, used to describe the size characteristics of the annotation box.

[0118] Cluster center / cluster center: the center point calculated from the data points by K-means++ clustering algorithm, as the anchor point of YOLO model.

[0119] Anchor point: the final result of cluster center, used in YOLO model, to help the model better predict the boundary box of target object.

[0120] The result of clustering: the final 9 cluster centers as the anchor points of YOLO model, which can better match the width and height ratio of the annotation box, thus improving the detection accuracy.

[0121] The purpose of the target detection step is to identify the live window area and poster recommendation area in the set-top box page through the improved YOLO model, and generate the boundary box.

[0122] Figure 3 The improved YOLO model in this application is shown in the results of set-top box page element detection, where the poster recommendation area is detected by a yellow annotation box, and the live window area is detected by a red annotation box. The beneficial effects of this step are:

[0123] 1. Automation and optimization of anchor point selection: compared with the fixed anchor box obtained by coco dataset in traditional YOLO model, K-means++ clustering algorithm is adopted, which can dynamically select suitable anchor points according to data distribution. This method automatically adapts to the diversity and complexity of target area in set-top box page, improves the detection accuracy and stability.

[0124] 2. Strong adaptability: by collecting data from multiple operator pages, a customized dataset for set-top box page is constructed, ensuring the diversity of training data and ensuring that model training can cover the diversity and complexity of different operator pages, so that the model has practical scene adaptability.

[0125] II. Fine-tuning PaddleOCR character recognition.

[0126] Based on the fine-tuned PaddleOCR model, all characters in the set-top box page are recognized, and the specific steps are as follows:

[0127] 1. Obtain set-top box page data. Real set-top box page image data is obtained by pulling code streams from set-top box pages of different operators.

[0128] 2. Labeling using PaddleOCR labeling tool:

[0129] (1) Use the pre-training model of PaddleOCR and the labeling tool PPOCRLabel to automatically label the set-top box page data.

[0130] (2) Check the labeling situation one by one, manually modify the mislabeled and missed characters, and ensure the accuracy of the labeling results.

[0131] (3) Assign a specific rotation angle to each text area technical domain to ensure that it can adapt to text at different inclination angles.

[0132] 3. After completing data labeling, export the labeling results, and divide the labeled images and corresponding labels into a second training set, a second validation set, and a second test set according to an 8:1:1 ratio.

[0133] 4. Download the detection model, recognition model, and direction classification model of ch_ppocr_server_v4.0 and their respective configuration files, and fine-tune the model based on them.

[0134] 5. Fine-tune the model, train the detection model, the recognition model, and the direction classifier. In training the detection model, the hyperparameters are as follows: algorithm model: DB++, loss function: BD++Loss, optimizer: Adam, learning rate: 0.001, and detection box threshold: 0.6. In training the recognition model, the hyperparameters are as follows: algorithm model: SVTR_HGNet, loss function: MultiLoss, optimizer: Adam, and learning rate: 0.001. In training the direction classifier, the hyperparameters are as follows: algorithm model: MobileNetV3, loss function: ClsLoss, angle label list: ['0', '180', '45', '-45'], optimizer: Adam, and learning rate: 0.001. These training hyperparameters are optimized through experiments and can provide good results on the set-top box page dataset. (Use the second training set to train the detection model, the recognition model, and the direction classification model, and use the second validation set to verify the performance of each model, and adjust the corresponding hyperparameters of each model according to the verification results.)

[0135] 6. After training is completed, test and evaluate on the second test set. The results show that the fine-tuned model can effectively improve the detection and recognition ability on set-top box page data without affecting the general detection ability. After evaluation is completed, the second deep learning model is obtained.

[0136] Figure 4 The fine-tuned PaddleOCR model demonstrates the results of character recognition in the set-top box page. All characters can be seen, including the artistic fonts in the poster recommendation position, which can also be accurately detected and recognized. The benefits of this step are:

[0137] 1. Artistic font recognition: The fine-tuned model can effectively recognize artistic fonts in the set-top box page. These fonts are often missed by traditional OCR models due to their irregular shape or unique design.

[0138] 2. Inclined text recognition: The PaddleOCR model's built-in text direction classifier only supports 0-degree and 180-degree rotation classification. However, there are texts with 45-degree and -45-degree angles in the set-top box page. By fine-tuning the angle classifier, the model can accurately recognize texts with different rotation angles, solving the common problem of text inclination in the set-top box page.

[0139] In the technical solution of the present application, the second deep learning model used in the character recognition step (i.e., the fine-tuned PaddleOCR model) can accurately recognize characters in the set-top box page image and generate multiple text boxes. The model includes a detection model, a recognition model, and a text direction classification model, which work together to achieve efficient end-to-end character recognition. Each text box not only contains the recognized text content but also accurately identifies the position coordinates (bounding box information) of the text region in the image. This makes the character recognition result not only obtain the text content but also obtain the spatial position information closely related to the text content, laying a solid foundation for subsequent page element analysis and layout pattern recognition.

[0140] The detection model of the PaddleOCR model can finely capture the visual spatial distribution and independence of the text region in the image. For example, for "Hunan Satellite TV" and "Jiangsu Satellite TV" arranged vertically in the same column, since they are separate text lines in visual space, the detection model will generate independent text boxes for them. This mechanism effectively separates text that is spatially separated and usually represents independent semantic units (such as different channel names). This text box generation result based on visual spatial independence more accurately reflects the actual layout and content structure of the page, providing more structurally meaningful and semantically related input for subsequent element merging and clustering grouping.

[0141] By optimizing the hyperparameters of the detection model, recognition model, and direction classification model, and combining training data with rich styles, the recognition robustness of special fonts is significantly enhanced, including accurate recognition of oblique fonts and artistic fonts, which is often ineffective in traditional OCR models. The synergistic optimization of the three sub-models ensures the dual improvement of efficiency and accuracy in the entire character recognition process.

[0142] III. Element merging

[0143] In the set-top box page, there may be overlapping or adjacent situations between text boxes and image boxes (such as live window, poster recommendation position, etc.). In order to improve the accuracy of element recognition, it is necessary to merge these overlapping or adjacent elements. The specific steps are as follows:

[0144] 1. Traverse all detected elements and store text boxes and image boxes separately. Text boxes are obtained through the character recognition module, while image boxes are obtained through the target detection module.

[0145] 2. Text box merging:

[0146] (1) Sort the text boxes by the horizontal coordinate of the top-left corner.

[0147] (2) Traverse each text box and check adjacent or overlapping text boxes.

[0148] (3) Calculate the intersection area of two text boxes.

[0149] (4) If the intersection area of two text boxes is greater than 0, merge them into one text box. If there are more than two text boxes with intersection area greater than 0, only merge the two text boxes with the largest intersection area. The merged text box no longer participates in the subsequent merging process.

[0150] (5) When merging, update the position and text content of the merged text box. The updated position is the smallest closed region containing both text boxes, and the text content is the text content of the larger text box. Record the removed text box content as an auxiliary attribute of the preserved text content.

[0151] 3. Text box distance merging:

[0152] (1) For the remaining unmerged text boxes, check the distance of adjacent text boxes.

[0153] (2) Calculate the horizontal distance and vertical distance between two text boxes. If the horizontal distance or vertical distance between two text boxes is less than 10 pixels, merge them into one text box. If there are more than two text boxes with a horizontal distance or vertical distance less than 10 pixels, only merge the two text boxes with the smallest distance. The merged text box no longer participates in subsequent merging.

[0154] (3) When merging, update the position and text content of the merged text box. The updated position is the smallest closed area containing both text boxes, and the text content is the text content of the larger text box. Record the removed text box content as an auxiliary attribute of the preserved text content.

[0155] 4. Merge text boxes with image boxes:

[0156] (1) For the remaining text boxes that do not participate in merging, check if they intersect with image boxes.

[0157] (2) If the intersection area of the text box and the image box is greater than 0, remove such text boxes and record the removed text box content as an auxiliary attribute of the image box.

[0158] 5. Check the inclusion relationship between elements and establish master-slave relationship. The specific master-slave relationship is as follows:

[0159] (1) During the merging of text boxes, the preserved text box is the master element and the removed text box is the auxiliary element.

[0160] (2) During the merging of text boxes and image boxes, the image box is the master element and the removed text box in the image box is the auxiliary element.

[0161] 6. For each preserved master element, when judging its charging attribute, it can be determined whether the element contains "free", "paid", "VIP", "SVIP" and other keywords according to the content of the removed text box, and the charging attribute is updated accordingly, that is, the charging attribute of the master element is determined according to the content of the auxiliary element; wherein the charging attribute is one of the above-mentioned auxiliary attributes.

[0162] Figure 5 The merging result of the text box and the image box is shown, and it can be seen that the overlapping and closely spaced text boxes in the original result have been successfully merged, and the text boxes in the poster recommendation position have been removed. This step has the following beneficial effects:

[0163] 1. Intelligent merging of text boxes: During the merging process, the distance and relative position between text boxes are judged to ensure that only truly close text boxes are merged, avoiding false merging.

[0164] 2. Intelligent merging of text boxes and image boxes: By removing text boxes and attaching their contents to image boxes, duplicate identification of text boxes and image elements is avoided, improving the integration and identification accuracy of page elements.

[0165] Four, element optimization.

[0166] After completing the element merging, further optimization of the elements is needed to improve the accuracy of page layout analysis. The specific steps are as follows:

[0167] 1. In the character recognition results, some text boxes may contain meaningless noise. Remove text boxes with a text length less than or equal to 1 and some meaningless identifier symbols.

[0168] 2. To avoid subsequent processing difficulties caused by duplicate or disordered identifiers, reassign element IDs to text boxes and image boxes to ensure that IDs start from 1 and gradually increase, avoiding empty or duplicate IDs.

[0169] The beneficial effects of this step are: efficient filtering of text: through length and rule double filtering, meaningless text can be effectively removed, reducing noise interference on subsequent analysis.

[0170] Five, clustering and grouping.

[0171] After completing the element merging and optimization, the elements need to be clustered and grouped to identify different element groups in the page, such as live window, status bar, navigation bar, poster recommendation position, etc. Clustering and grouping can help identify the relationship and layout pattern between elements. The specific steps are as follows:

[0172] 1. Load the processed element data, which has been processed through target detection, character recognition, element merging and optimization.

[0173] 2. Use the DBSCAN algorithm to cluster elements based on their size and position attributes (width, height, position) to identify element groups with similar sizes. In the DBSCAN algorithm, "unvisited points" refer to: all initial points (all elements processed by element optimization) are unvisited points, that is, they have not been checked for core points, nor have they been grouped into a cluster or marked as noise points. The actual meaning of the point here is the coordinate position of the merged element. The specific method is as follows:

[0174] (1) Select an unvisited point P from the data set. (In the first round of clustering, since all points (elements) are unvisited points, a point is randomly selected as the current processing point P; in subsequent clustering, a point is randomly selected from the remaining unvisited points).

[0175] (2) Check the epsilon neighborhood of the point, count the number of points contained in the neighborhood. If the number of points in the neighborhood is greater than or equal to the minimum number of border points, then the point is a core point.

[0176] (3) Extend the cluster, add all points in the neighborhood of the core point to the cluster, for each point in the neighborhood, if it is also a core point, continue to extend the neighborhood of the point, recursively until no new core point can be extended.

[0177] (4) Handle noise, if a point is not a core point and does not belong to the neighborhood of any core point, mark it as noise.

[0178] (5) Repeat continuously, perform the above operations on all points in the dataset until all points are visited.

[0179] After multiple attempts to optimize, the best parameters for DBSCAN are epsilon = 30 and min_samples = 3.

[0180] 3. Modify the preliminary clustering results based on the size and position of the elements, check if adjacent elements belong to the same group, if there is a significant connection across groups, merge these groups, ensure that the distance or size change between elements in the same group is not too large, if significant inconsistencies are found, adjust the clustering division. Among them, DBSCAN is a density-based clustering algorithm, the initial clustering result may appear broken or missing, causing elements that should be in the same group to be clustered into different groups, which is a significant connection across groups. For such cross-group results, merging is required. For example Figure 6 In the navigation bar group above, it is possible that during clustering, the middle element assembly is broken into two or even three groups.

[0181] 4. Group elements according to the clustering results and further record the alignment of each element group. Analyze the relative position relationship of the element group, identify which elements are horizontally arranged and which are vertically arranged, and ensure the accuracy of the layout relationship.

[0182] 5. An element may be included in multiple clustering results, causing clustering conflicts. To solve such conflicts, elements can be assigned to the most suitable group according to the following two methods:

[0183] (1) Average area judgment: calculate the average area of the element, assign the element to the group with larger area, and ensure that each element is reasonably classified.

[0184] (2) Relative position adjustment: if the element is located at the intersection of multiple clusters, determine its final attribution based on its relative position, size, and other attributes.

[0185] 6. By comparing the layout features of multiple element groups, layout patterns with similar structures are identified.

[0186] 7. For isolated elements that did not participate in the initial clustering, further analysis is needed to re-group them. The spatial gap between un-clustered elements and clustered elements is analyzed, and if they are close in distance or have a clear relative position relationship, the un-clustered elements are re-assigned to the corresponding group.

[0187] Figure 6 It is shown how the DBSCAN algorithm is used to cluster the identified elements to identify different element groups in the set-top box page (status bar group, navigation bar group, recommendation position group). The benefits of this step are:

[0188] 1. Improve page layout analysis accuracy: By accurately clustering and optimizing different types of elements in the set-top box page (such as live window, poster recommendation position, text box, etc.), the system can identify the relationship and hierarchy between different element groups.

[0189] 2. Automatically adapt to diverse page layouts: The layout of the set-top box page is usually diverse and complex, especially the design styles of different operators. By using the DBSCAN clustering algorithm, the system can automatically adapt to these different layout patterns, whether it is horizontal or vertical arrangement, or even complex cross layout, clustering and grouping can be flexibly handled.

[0190] This application combines traditional vision and deep learning technology to propose an efficient and accurate set-top box page element detection and layout analysis method. This method can significantly improve the recognition accuracy and layout analysis capability of set-top box page elements, with the following specific effects:

[0191] 1. Improve recognition accuracy: Through the improved YOLO model and fine-tuned PaddleOCR model, this application can effectively identify various elements in the set-top box page, especially for complex text (such as artistic fonts, slanted fonts, etc.) and diverse layouts, with significantly improved recognition accuracy.

[0192] 2. Effective information fusion: This method combines object detection and character recognition technology, not only can identify objects in the set-top box page (such as live window, poster recommendation position, etc.), but also can accurately extract text information and effectively fuse it with image elements, solving the problem that the target detection and character recognition results cannot be effectively combined in the prior art.

[0193] 3. Hierarchical layout analysis: Through multi-level analysis of elements, the application can provide more comprehensive information for page layout, helping to better understand the structure and design rules of the set-top box page. This not only helps page interaction and design optimization, but also provides support for subsequent user experience analysis.

[0194] 4. Wide applicability: The technology is not only suitable for set-top box page element detection and layout analysis, but also can be extended to other complex user interface recognition and optimization tasks, with high universality and application prospects.

[0195] The technology proposed in this application has been tested in our company's set-top box picture recognition engine and has achieved good implementation results. Through efficient and accurate picture recognition and layout analysis capabilities, it can monitor and identify the picture content output by the set-top box in real time, ensuring the compliance and quality of the content, and achieving good implementation results.

[0196] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described, but as long as the combinations of the technical features do not exist, they should be considered as the scope of the present application.

Claims

1.A deep learning-based page layout and element analysis method, characterized by, Comprise: Target detection: based on the first deep learning model to set top box page image target detection, identify the image element area in the page, and generate a plurality of image frames, according to the area size of the image frame to determine the live window area and poster recommendation position area: the largest image frame is the live window area, and the remaining image frame is the poster recommendation position area; Character recognition: based on the second deep learning model to recognize the characters in the set top box page image, generate a plurality of text frames; the first deep learning model and the second deep learning model are trained by pulling the code stream from the set top box page of different operators to obtain the real set top box page image to construct the data set; Element merging: each image frame and text frame is regarded as a separate original element, the overlapping condition, distance relationship between text frames and the overlapping condition between text frames and image frames are judged, and the overlapping or adjacent original elements are merged, wherein one original element is regarded as the main element and the other original elements are regarded as the subsidiary elements; For the merged elements, the subsidiary attributes of the main elements are determined according to the contents of the subsidiary elements; Clustering grouping: according to the size and position characteristics of the merged elements, clustering processing is carried out, and according to the clustering result, the elements are grouped, including status bar group, navigation bar group and poster recommendation position group; the alignment mode of each element group is recorded, the relative position relationship of the element groups is analyzed, and the layout mode of the element page is determined. 2.The deep learning-based page layout and element analysis method of claim 1, wherein, The first deep learning model is an improved YOLO model, specifically using K-means++ clustering algorithm to dynamically select suitable anchor points according to data distribution. 3.The deep learning-based page layout and element analysis method of claim 2, wherein, The training process of the first deep learning model comprises: S11, using image labeling tool LabelImg to label the live window area and poster recommendation position area of each image in the set top box page image data set, using the minimum circumscribed rectangle method to label the boundary area of each target, and the label is rectangle; S12, obtaining the real boundary size of all labeled frames from the label file derived after labeling, and obtaining the width and height value of all labeled frames; S13, randomly dividing the image data and label data into a first training set, a first validation set and a first test set according to a first preset ratio; S14, using K-means++ to cluster the labeled frames to obtain new anchor points; S15, replacing the generated anchor point data into the configuration file of YOLO model to obtain the improved YOLO model, training the improved YOLO model using the first training set, verifying the model performance using the first validation set during the training process; after the model training and verification are completed, the performance of the model is evaluated using the first test set, and the first deep learning model is obtained. 4.The deep learning-based page layout and element analysis method of claim 3, wherein, The step S14 specifically comprises: S141, selecting an initial clustering center: randomly selecting a data point from all labeled frames as the first clustering center; S142, calculating the Euclidean distance: for each remaining data point, calculating the Euclidean distance between it and the nearest selected clustering center, specifically: where d(x i , c i ) denotes the Euclidean distance, k denotes the dimension index, n denotes the dimension of the space, x i denotes the data point, c i denotes the cluster center; x ik denotes the data point in k-dimension; c ik denotes the cluster center in k-dimension; S143, update cluster center: select the next cluster center in a probabilistic way according to the distance of each data point to the nearest cluster center, select data point x i The probability of being the next cluster center is proportional to the square of its distance, and the specific formula is: where P(x i ) represents the probability of selecting data point x i , d(x i , c i ) represents the distance of data point x i to the nearest cluster center c i , X represents the set of all data points, x j represents the jth data point, and d(x j, c i ) represents the distance of the jth data point to the nearest cluster center c i ; S144, repeating step S132 and step S133 until a plurality of cluster centers are selected; S145, performing standard K-means clustering, assigning each data point to the nearest cluster center, and then updating the cluster centers until the clustering result converges, obtaining the final plurality of cluster centers as anchor points; wherein the number of cluster centers is determined according to the needs of the model. 5.The deep learning-based page layout and element analysis method of claim 1, wherein, The second deep learning model enhances the recognition ability of special fonts by optimizing hyperparameters. 6.The deep learning-based page layout and element analysis method of claim 5, wherein, The training process of the second deep learning model specifically includes: S21, using PaddleOCR labeling tool to label the boundary region of characters in the set-top box page image data, including: S211, using the pre-training model of PaddleOCR and the labeling tool PPOCRLabel to automatically label the character region of the set-top box page image data, and initially generating a labeling result; S212, checking the labeling situation one by one, and manually modifying the mislabeled or missed characters; S213, specifying the specific rotation angle for each text region to ensure that it can adapt to text with different inclination angles; S22, after completing the labeling of the character region, exporting the labeling result, and dividing the labeled image and the corresponding label into a second training set, a second validation set and a second test set according to a second preset proportion; S23, downloading the detection model, recognition model and direction classification model of the PaddleOCR model and their respective configuration files, and performing model fine-tuning on this basis; S24, model fine-tuning: using the second training set to train the detection model, recognition model and direction classification model of the PaddleOCR model, simultaneously verifying the performance of each model using the second validation set, and adjusting the hyperparameters corresponding to each model according to the verification result; S25, after the completion of model training and verification, using the second test set to evaluate the performance of the model, obtaining the second deep learning model. 7.The deep learning-based page layout and element analysis method of claim 1, wherein, The element merging step specifically includes: S31, traversing all detected elements, and storing text boxes and image boxes separately; wherein the text boxes are obtained through the character recognition step, and the image boxes are obtained through the target detection step; S32, text box merging, including: S321, sorting the text boxes according to the horizontal coordinate of the upper left corner; S322, traversing each text box and checking the text boxes adjacent or overlapping with it; S323, calculating the intersection area of the text box being checked and other text boxes adjacent or overlapping with it; S324, if the intersection area of the two text boxes is greater than 0, merging them into one text box, if there are multiple text boxes with intersection area greater than 0, only merging the two text boxes with the largest intersection area, and the merged text box no longer participates in the subsequent merging process; S325, when merging, updating the position and text content of the merged text box; the updated position is the smallest closed region containing both text boxes, and the text content is the text content of the text box with larger area, and the removed text box content is recorded as an attached attribute of the preserved text content; S33, merging of text boxes, including: S331, for the remaining unmerged text boxes, checking the distance to the adjacent text boxes; S332, calculating the horizontal distance and vertical distance between the current text box and the adjacent text boxes, if the horizontal distance or vertical distance between two text boxes is less than a preset number of pixels, merging them into one text box; if the horizontal distance or vertical distance between more than two text boxes is less than a preset number of pixels, only merging the two text boxes with the smallest distance; the merged text box no longer participates in the subsequent merging process; S333, when merging, updating the position and text content of the merged text box; the updated position is the smallest closed region containing both text boxes, and the text content is the text content of the text box with larger area, and the removed text box content is recorded as an attached attribute of the preserved text content; S34, merging of text boxes and image boxes, including: S341, for the remaining text boxes that have not participated in merging, checking whether they have intersection with image boxes; S342, if the intersection area of the text box and the image box is greater than 0, then remove such text box, and record the removed text box content as an attached attribute of the image box. 8.The deep learning-based page layout and element analysis method of claim 1, wherein, The attached attribute of the main element includes a charging attribute; the content of the attached element includes a keyword about whether the user needs to pay, and the charging attribute of the main element is determined according to the keyword. 9.The deep learning-based page layout and element analysis method of claim 1, wherein, Before clustering grouping, element optimization is also performed, including: S41, removing text boxes with text length less than or equal to 1 and meaningless identifier; S42, for the merged elements, reassigning element IDs to ensure that element IDs start from 1 and gradually increase, avoiding empty IDs or duplicate IDs. 10.The deep learning-based page layout and element analysis method of claim 1, wherein, The step of clustering grouping specifically includes: S51, using DBSCAN algorithm to cluster according to the size and position characteristics of the merged elements to identify element groups with similar sizes, specifically including: S511, randomly selecting an unvisited element from the elements as the current processing point P of the DBSCAN algorithm; S512, checking the ε neighborhood of the point P and calculating the number of unvisited points contained in the neighborhood except the point P, if the number of points in the neighborhood is greater than or equal to the minimum boundary point number of the ε neighborhood, then the point P is taken as a core point; S513, adding all points in the neighborhood of the core point to the cluster, for each point in the neighborhood, if it is also a core point, continue to expand the neighborhood of the point, and recursively until no new core point can be expanded; S514, processing noise: if a point is not a core point and does not belong to the neighborhood of any core point, mark it as noise and remove it; S515, performing steps S511-S514 on all points until all points are visited; S52, correct the initial clustering results, check whether the adjacent elements belong to the same group based on the size and position characteristics of the elements, if there is obvious connection across the group, merge these groups, ensure that the distance or size change between the elements in the same group will not be too large, if significant inconsistency is found, adjust the clustering division; S53, group the elements according to the clustering results, including status bar group, navigation bar group, poster recommendation position group, and further record the alignment mode of each element group; analyze the relative position relationship of the element group, identify which elements are horizontally arranged and which elements are vertically arranged; S54, reassign the elements that cause clustering conflicts, the assignment method includes: S541, average area judgment: calculate the average area of the element, and assign the element to the group with larger area; S542, relative position adjustment: if the element is located at the junction of multiple clusters, determine its final attribution through the relative position and size attribute of the element; S55, by comparing the layout characteristics of multiple element groups, identify the layout mode with similar structure; S56, for isolated elements not participating in the initial clustering, analyze the spatial gap between the unclustered elements and the clustered elements, if the distance between them is close or there is obvious relative position relationship, reassign the unclustered element to the corresponding group.