Document title hierarchy determination method and device and storage medium
By extracting and utilizing the basic features and optimizing feature vectors of document titles, the document title and its level are solved, and the problem of high rules-making requirements and difficult to improve accuracy in the prior art is achieved, and the accuracy of document title recognition and method universality is achieved.
Patent Information
- Application Number
- CN202510218625.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-17
AI Technical Summary
When identifying document titles and determining their levels, existing rules-based methods have problems such as high rule-making requirements, many rules conflicts, and difficult to improve accuracy. They are not universal, have high development costs, and have a great impact on the accuracy of electronic documents in irregular formats.
By obtaining the text lines of the target document, extracting the basic feature vectors and optimizing feature vectors, and using these feature vectors to determine the document title and its hierarchy. Optimized features include font application features, bold features, serial number features, font size features, etc., and the pre-trained classification model is used to set bold features.
It improves the accuracy of identification of document titles, reduces the requirements for rule formulation, enhances the universality of the method, reduces development costs, and can effectively deal with electronic documents in irregular formats.
Smart Images

Figure CN120163146A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a method, apparatus, and storage medium for determining the hierarchy of document titles. Background Art
[0002] Electronic documents, such as PDF documents, Word documents, RTF (Rich Text Format) documents, and HTML (HyperText Markup Language) documents, etc., are the main media forms for carrying information in various computer systems and are widely used. Therefore, extracting valuable information from electronic documents has gradually become a research hotspot in recent years. And document titles can help users quickly obtain the core content of the document.
[0003] In the related art, to extract the document title from an electronic document and determine the hierarchy of the document title, the commonly used method is to identify the document title based on rules. This method formulates some extraction rules for the document title according to the difference between the text style of the document title and the text style of the body text, and uses these rules to extract the document title from the electronic document and determine the hierarchy of the document title.
[0004] However, this rule-based method has high requirements for rule formulation, and conflicts are prone to occur between rules, resulting in difficulty in improving the recognition accuracy of document titles. In addition, the rule-based method is not universal. When the text styles of different electronic documents are diverse, extraction rules must be formulated separately, resulting in a high development cost. Moreover, the non-standard formats of some electronic documents will also affect the accuracy of this rule-based method. Summary of the Invention
[0005] To solve the technical problems that the above-mentioned rule-based method has high requirements for rule formulation, and conflicts are prone to occur between rules, resulting in difficulty in improving the recognition accuracy of document titles. In addition, the rule-based method is not universal. When the text styles of different electronic documents are diverse, extraction rules must be formulated separately, resulting in a high development cost. Moreover, the non-standard formats of some electronic documents will also affect the accuracy of this rule-based method, the embodiments of this application provide a method, apparatus, electronic device, and storage medium for determining the hierarchy of document titles. The specific technical solutions are as follows:
[0006] In the first aspect of the embodiments of this application, first, a method for determining the hierarchy of document titles is provided, and the method includes:
[0007] Obtain a target document, and obtain the text lines in the target document page of the target document, and store them in a title candidate line set;
[0008] Extract the basic feature vectors and optimized feature vectors for each text line in the set of title candidate lines;
[0009] Determine the document title in the set of title candidate lines according to the basic feature vectors and the optimized feature vectors, and determine the title level of the document title.
[0010] In an optional implementation manner, extracting the optimized feature vectors for each text line in the set of title candidate lines includes:
[0011] Extract the optimized features for each text line in the set of title candidate lines, and form an optimized feature vector from the optimized features;
[0012] Wherein, the optimized features include at least one of the following: font application feature, bold feature, serial number feature, font size feature, document boundary feature, page index feature, line index feature, text feature, color feature.
[0013] In an optional implementation manner, when the optimized feature is the font application feature, extracting the optimized features for each text line in the set of title candidate lines includes:
[0014] For any text line in the set of title candidate lines, obtain the fonts corresponding to each character in the text line;
[0015] When all the fonts are the same, set the font application feature to a first value;
[0016] When all or some of the fonts are different, set the font application feature to a second value.
[0017] In an optional implementation manner, when the optimized feature is the bold feature, extracting the optimized features for each text line in the set of title candidate lines includes:
[0018] For any text line in the set of title candidate lines, obtain the fonts corresponding to each character in the text line; when all the fonts are the same and are bold fonts, set the bold feature to a first value;
[0019] Or,
[0020] For any text line in the set of title candidate lines, obtain the fonts corresponding to each character in the text line and the font weights; when all the fonts are the same and all the font weights are greater than a preset weight threshold, set the bold feature to a first value;
[0021] Or,
[0022] For any text line in the set of title candidate lines, obtain the font and font weight corresponding to each character in the text line; in the case where all or some of the fonts are different, all the font weights are not greater than a preset weight threshold, and all the fonts are not bold fonts, perform the following processing; classify the text included in the text line according to the font types included in the text line, and set bold characteristics for the text included in the text line for different font types; wherein, classifying the text included in the text line according to the font types included in the text line and setting bold characteristics for the text included in the text line for different font types includes: rendering the text line by font to obtain at least two third text line images, and searching for the next text line of the text line from the set of title candidate lines; rendering the next text line to obtain a second text line image, and setting the bold characteristics according to at least two of the third text line images and the second text line image.
[0023] In an optional implementation manner, the method further includes:
[0024] In the case where all the fonts are the same, all the font weights are not greater than a preset weight threshold, and all the fonts are not bold fonts, perform the following processing;
[0025] Render the text line to obtain a first text line image, and search for the next text line of the text line from the set of title candidate lines;
[0026] Render the next text line to obtain a second text line image, and set the bold characteristics according to the first text line image and the second text line image.
[0027] In an optional implementation manner, the setting the bold characteristics according to the first text line image and the second text line image includes:
[0028] Input the first text line image and the second text line image into a pre-trained bold classification model to obtain a first output result;
[0029] In the case where the first output result is a preset result, set the bold characteristic to a first value;
[0030] Or, in the case where the first output result is not a preset result, set the bold characteristic to a second value.
[0031] In an optional implementation manner, the setting the bold characteristics according to at least two of the third text line images and the second text line image includes:
[0032] For any of the third text line images, input the third text line image and the second text line image into a pre-trained bold classification model to obtain a second output result;
[0033] When all the second output results are preset results, set the bold feature to a first value;
[0034] Alternatively, when all or part of the second output results are not preset results, set the bold feature to a second value.
[0035] In an alternative embodiment, the pre-trained bold classification model is obtained by the following method:
[0036] Obtain character image pairs and sample labels corresponding to the character image pairs, input the character image pairs into a bold classification model to obtain a prediction result;
[0037] Determine the loss value between the prediction result and the sample label, and train the bold classification model according to the loss value;
[0038] When the loss value converges, stop training to obtain a pre-trained bold classification model.
[0039] In an alternative embodiment, when the optimization feature is a serial number feature, the extraction of the optimization feature of each text line in the title candidate line set includes:
[0040] For any text line in the title candidate line set, detect whether there is a serial number character at the beginning of the text line; if there is a serial number character, find the serial number value corresponding to the serial number character, and set the serial number feature to the serial number value; or, if there is no serial number character, set the serial number feature to a second value;
[0041] Alternatively, when the optimization feature is a font size feature, the extraction of the optimization feature of each text line in the title candidate line set includes: for any text line in the title candidate line set, obtain the font sizes corresponding to each character in the text line; determine the average font size of the font sizes, and set the font size feature to the average font size;
[0042] Alternatively, when the optimization feature is a document boundary feature, the extraction of the optimization feature of each text line in the title candidate line set includes: for any text line in the title candidate line set, determine the distance between the text line and the document boundary, and set the document boundary feature to the distance;
[0043] Alternatively, when the optimization feature is a page index feature or a line index feature, extracting the optimization feature of each text line in the candidate title line set includes:
[0044] For any text line in the candidate title line set, determine the page index and line index of the text line; set the page index feature to the page index and the line index feature to the line index;
[0045] Alternatively, when the optimization feature is a text feature, extracting the optimization feature of each text line in the candidate title line set includes: for any text line in the candidate title line set, obtain the characters included in the text line, and set the text feature to the characters included in the text line;
[0046] Alternatively, when the optimization feature is a color feature, extracting the optimization feature of each text line in the candidate title line set includes: for any text line in the candidate title line set, determine the character region in the text line; determine the pixel average value of the character region, and set the color feature to the pixel average value.
[0047] In an alternative embodiment, determining the document title in the candidate title line set according to the basic feature vector and the optimization feature vector includes:
[0048] For any text line in the candidate title line set, input the basic feature vector and the optimization feature vector of the text line into a pre-trained title classifier to obtain a classification probability value;
[0049] When the classification probability value is greater than a preset classification threshold, determine the text line as the document title, and store the optimization feature vector of the text line in the title feature set;
[0050] Alternatively, when the classification probability value is not greater than the preset classification threshold, store the optimization feature vectors of the text lines that meet the first preset condition in the candidate title feature set; determine the title level of the document title according to the title feature set and the candidate title feature set.
[0051] In an alternative embodiment, the pre-trained title classifier is obtained by the following method:
[0052] Obtain a sample title corpus, and extract the sample basic characteristic vector and the sample optimization feature vector of the sample title corpus;
[0053] Input the sample basic characteristic vector and the sample optimization feature vector into a title classifier to obtain a predicted classification probability value;
[0054] Determine the probability loss between the predicted classification probability value and the sample classification probability value corresponding to the sample title corpus;
[0055] Train the title classifier according to the probability loss, and stop training when the probability loss converges to obtain a pre-trained title classifier.
[0056] In an alternative embodiment, the determining the title level of the document title includes:
[0057] Divide the title feature set into a first title feature set and a second title feature set;
[0058] Divide the candidate title feature set into a first candidate title feature set and a second candidate title feature set;
[0059] Determine the first candidate optimization feature vector in the first candidate title feature set, store it in the first title feature set to obtain a first target title feature set, and determine the title level of the text line corresponding to the optimization feature vector in the first target title feature set;
[0060] Determine the second candidate optimization feature vector in the second candidate title feature set, store it in the second title feature set to obtain a second target title feature set, and determine the title level of the text line corresponding to the optimization feature vector in the second target title feature set.
[0061] In an alternative embodiment, the dividing the title feature set into a first title feature set and a second title feature set includes:
[0062] Traverse the optimization feature vectors of the title feature set to obtain the serial number feature in the traversed optimization feature vector;
[0063] If the serial number feature is not the second value, store the optimization feature vector in the first title feature set;
[0064] If the serial number feature is the second value, store the optimization feature vector in the second title feature set;
[0065] The dividing the candidate title feature set into a first candidate title feature set and a second candidate title feature set includes:
[0066] Traverse the optimization feature vectors in the candidate title feature set to obtain the serial number feature in the traversed optimization feature vector;
[0067] When the serial number feature is not the second numerical value, store the optimized feature vector in the first candidate title feature set;
[0068] When the serial number feature is the second numerical value, store the optimized feature vector in the second candidate title feature set.
[0069] In an optional implementation manner, determining the first candidate optimized feature vector in the first candidate title feature set and storing it in the first title feature set to obtain the first target title feature set, and determining the title level of the text line corresponding to the optimized feature vector in the first target title feature set includes:
[0070] Cluster the optimized feature vectors in the first title feature set to obtain multiple first clustering sets, and determine the benchmark feature vector corresponding to each first clustering set;
[0071] Traverse the optimized feature vectors in the first candidate title feature set, and determine the first similarity between the traversed optimized feature vector and any one of the benchmark feature vectors;
[0072] Find the largest first target similarity from the first similarities, and determine whether the first target similarity is greater than a preset first similarity threshold;
[0073] When the first target similarity is greater than the preset first similarity threshold, determine the traversed optimized feature vector as the first candidate optimized feature vector;
[0074] Store the first candidate optimized feature vector in the first title feature set to obtain the first target title feature set, and determine the title level of the text line corresponding to the optimized feature vector in the first target title feature set;
[0075] Among them, storing the first candidate optimized feature vector in the first title feature set to obtain the first target title feature set, and determining the title level of the text line corresponding to the optimized feature vector in the first target title feature set includes:
[0076] Determine the benchmark feature vector corresponding to the first target similarity, find the first clustering set corresponding to the determined benchmark feature vector, and store the first candidate optimized feature vector in the found first clustering set;
[0077] Determine the title level of the text line corresponding to the optimized feature vector in each updated first clustering set; wherein, each updated first clustering set is combined to obtain the first target title feature set.
[0078] In an alternative embodiment, clustering the optimized feature vectors in the first title feature set to obtain a plurality of first clustering sets includes:
[0079] Obtain the serial number features included in the optimized feature vectors in the first title feature set, and determine whether the serial number features are the same;
[0080] In the case where the serial number features are different, cluster the optimized feature vectors in the first title feature set according to the serial number features to obtain a plurality of first clustering sets;
[0081] Alternatively, in the case where the serial number features are the same, cluster the optimized feature vectors in the first title feature set according to the font size features to obtain a plurality of first clustering sets.
[0082] In an alternative embodiment, determining the reference feature vector corresponding to each of the first clustering sets includes:
[0083] For any one of the first clustering sets, extract the reference features corresponding to the first clustering set, and form a reference feature vector from the reference features;
[0084] Wherein, the reference features include at least one of the following: reference font application feature, reference bold feature, reference serial number feature, reference font size feature, reference document boundary feature, reference page index feature, reference line index feature, reference text feature, reference color feature.
[0085] In an alternative embodiment, when the reference feature is the reference font application feature, extracting the reference features corresponding to the first clustering set includes: counting the first font application quantity corresponding to the optimized feature vectors with the font application feature being the first value in the first clustering set; counting the second font application quantity corresponding to the optimized feature vectors with the font application feature being the second value in the first clustering set; in the case where the first font application quantity is greater than the second font application quantity, setting the reference font application feature corresponding to the first clustering set to the first value; or, in the case where the first font application quantity is not greater than the second font application quantity, setting the reference font application feature corresponding to the first clustering set to the second value;
[0086] Alternatively, when the reference feature is a reference bold feature, the extraction of the reference feature corresponding to the first clustering set includes: counting the first bold quantity corresponding to the optimized feature vectors with the bold feature being the first value in the first clustering set; counting the second non-bold quantity corresponding to the optimized feature vectors with the bold feature being the second value in the first clustering set; when the first bold quantity is greater than the second non-bold quantity, setting the reference bold feature corresponding to the first clustering set to the first value; or, when the first bold quantity is not greater than the second non-bold quantity, setting the reference bold feature corresponding to the first clustering set to the second value;
[0087] Alternatively, when the reference feature is a reference serial number feature, the extraction of the reference feature corresponding to the first clustering set includes: obtaining the serial number feature included in the optimized feature vectors in the first clustering set, and setting the reference serial number feature corresponding to the first clustering set to the serial number feature;
[0088] Alternatively, when the reference feature is a reference font size feature, the extraction of the reference feature corresponding to the first clustering set includes: obtaining the font size feature included in the optimized feature vectors in the first clustering set, and determining the average font size feature of the font size feature; setting the reference font size feature corresponding to the first clustering set to the average font size feature;
[0089] Alternatively, when the reference feature is a reference document boundary feature, the extraction of the reference feature corresponding to the first clustering set includes: obtaining the document boundary feature included in the optimized feature vectors in the first clustering set, and determining the average document boundary feature of the document boundary feature; setting the reference document boundary feature corresponding to the first clustering set to the average document boundary feature.
[0090] Alternatively, when the reference feature is a reference page index feature or a reference line index feature, the extraction of the reference feature corresponding to the first clustering set includes: obtaining the page index feature included in the optimized feature vectors in the first clustering set, and selecting the smallest page index feature from the page index features; setting the reference page index feature corresponding to the first clustering set to the smallest page index feature; obtaining the line index feature included in the optimized feature vectors in the first clustering set, and selecting the smallest line index feature from the line index features; setting the reference line index feature corresponding to the first clustering set to the smallest line index feature;
[0091] Alternatively, when the reference feature is a reference text feature, the extraction of the reference feature corresponding to the first clustering set includes: setting the reference text feature corresponding to the first clustering set to be empty, where empty means not filling in any content;
[0092] Alternatively, when the reference feature is a reference color feature, the extracting the reference feature corresponding to the first clustering set includes: obtaining the color features included in the optimized feature vectors in the first clustering set, and determining the average color feature of the color features; setting the reference color feature corresponding to the first clustering set as the average color feature.
[0093] In an alternative embodiment, the determining the title levels of the text lines corresponding to the optimized feature vectors in each updated first clustering set includes:
[0094] Sorting all the reference feature vectors to obtain a first sorting result;
[0095] According to the first sorting result, determining the title levels of the text lines corresponding to the optimized feature vectors in each updated first clustering set;
[0096] Wherein, the determining the title levels of the text lines corresponding to the optimized feature vectors in each updated first clustering set according to the first sorting result includes:
[0097] Traversing the reference feature vectors in the first sorting result in order:
[0098] When the traversed reference feature vector is the first reference feature vector, searching for the updated first clustering set corresponding to the traversed reference feature vector; determining the first-level title corresponding to the traversed reference feature vector as the title level of the text line corresponding to the optimized feature vector in the updated first clustering set corresponding to the traversed reference feature vector;
[0099] Alternatively, when the traversed reference feature vector is not the first reference feature vector, searching for the previous reference feature vector of the traversed reference feature vector from the first sorting result; determining the title level of the text line corresponding to the optimized feature vector in the updated first clustering set corresponding to the traversed reference feature vector according to the previous reference feature vector.
[0100] In an alternative embodiment, the determining the title level of the text line corresponding to the optimized feature vector in the updated first clustering set corresponding to the traversed reference feature vector according to the previous reference feature vector includes:
[0101] When the reference page index feature included in the previous reference feature vector is the same as the reference page index feature included in the traversed reference feature vector, determining the previous title level corresponding to the previous reference feature vector;
[0102] Determining the sum of the previous title level and a preset level threshold as the title level corresponding to the traversed reference feature vector;
[0103] Determine the title level corresponding to the text line of the optimized feature vector in the updated first clustering set corresponding to the traversed reference feature vector as the title level corresponding to the traversed reference feature vector;
[0104] Add the traversed reference feature vector to the sub - title list of the previous reference feature vector;
[0105] Alternatively, in the case where the reference page index feature included in the previous reference feature vector is different from the reference page index feature included in the traversed reference feature vector, sort the optimized feature vectors in each updated first clustering set to obtain a second sorting result;
[0106] For the updated first clustering set corresponding to the traversed reference feature vector, traverse the optimized feature vectors in the updated first clustering set, and find the previous optimized feature vector of the traversed optimized feature vector from the second sorting result;
[0107] In the case where there is no previous optimized feature vector, find the next optimized feature vector of the traversed optimized feature vector from the second sorting result; or, in the case where there is a next optimized feature vector, find the reference feature vector corresponding to the updated first clustering set where the next optimized feature vector is located;
[0108] In the case where the traversed reference feature vector is different from the reference feature vector corresponding to the updated first clustering set where the next optimized feature vector is located, and the product of the reference font size feature included in the next optimized feature vector and a preset first multiple is less than the reference font size feature included in the traversed reference feature vector, determine the next title level of the reference feature vector corresponding to the updated first clustering set where the next optimized feature vector is located;
[0109] Determine the difference between the next title level and a preset level threshold as the title level corresponding to the traversed reference feature vector; determine the title level corresponding to the text line of the optimized feature vector in the updated first clustering set corresponding to the traversed reference feature vector as the title level corresponding to the traversed reference feature vector;
[0110] Add the reference feature vector corresponding to the updated first clustering set where the next optimized feature vector is located to the sub - title list of the traversed reference feature vector.
[0111] In an alternative embodiment, the method further includes:
[0112] In the case where there is a previous optimized feature vector, determine the previous title level of the reference feature vector corresponding to the updated first clustering set where the previous optimized feature vector is located;
[0113] Determine the title level corresponding to the benchmark feature vector traversed as the sum of the previous title level and the preset level threshold;
[0114] Determine the title level of the text line corresponding to the optimized feature vector in the updated first clustering set corresponding to the benchmark feature vector traversed as the title level corresponding to the benchmark feature vector traversed;
[0115] Add the benchmark feature vector traversed to the sub - title list of the subsequent benchmark feature vector.
[0116] In an alternative embodiment, determining the second candidate optimized feature vector in the second candidate title feature set and storing it in the second title feature set to obtain the second target title feature set, and determining the title level of the text line corresponding to the optimized feature vector in the second target title feature set includes:
[0117] Cluster the optimized feature vectors in the second title feature set to obtain multiple second clustering sets;
[0118] Traverse the optimized feature vectors in the second candidate title feature set, and determine the second similarity between the traversed optimized feature vector and the central optimized feature vector in any second clustering set;
[0119] Find the maximum second target similarity from the second similarities, and determine whether the maximum second target similarity is greater than the preset second similarity threshold;
[0120] In the case where the maximum second target similarity is greater than the preset second similarity threshold, determine the traversed optimized feature vector as the second candidate optimized feature vector;
[0121] Store the second candidate optimized feature vector in the second title feature set to obtain the second target title feature set, and determine the title level of the text line corresponding to the optimized feature vector in the second target title feature set;
[0122] Wherein, storing the second candidate optimized feature vector in the second title feature set to obtain the second target title feature set, and determining the title level of the text line corresponding to the optimized feature vector in the second target title feature set includes:
[0123] Find the second clustering set corresponding to the maximum second target similarity, and store the second candidate optimized feature vector in the found second clustering set;
[0124] Merge the updated second clustering sets to obtain the second target title feature set, and determine the title level of the text line corresponding to the optimized feature vector in the second target title feature set.
[0125] In an alternative embodiment, clustering the optimized feature vectors in the second set of title features to obtain a plurality of second clustering sets, including:
[0126] Clustering the optimized feature vectors in the second set of title features according to bold feature, font size feature, and document boundary feature to obtain a plurality of second clustering sets.
[0127] In an alternative embodiment, determining the title level of the text line corresponding to the optimized feature vector in the second set of target title features includes:
[0128] Merging the first set of target title features and the second set of target title features to obtain a third set of target title features;
[0129] Sorting the optimized feature vectors in the third set of target title features to obtain a third sorting result;
[0130] Traversing the optimized feature vectors in the second set of target title features and finding the position index corresponding to the traversed optimized feature vector in the third sorting result;
[0131] Determining the title level of the text line corresponding to the traversed optimized feature vector according to the position index.
[0132] In an alternative embodiment, determining the title level of the text line corresponding to the traversed optimized feature vector according to the position index includes:
[0133] Centering on the position index, forward searching in the third sorting result for the nearest neighbor forward optimized feature vector with the serial number feature being the serial number value, and backward searching for the nearest neighbor backward optimized feature vector with the serial number feature being the serial number value;
[0134] Determining the title level of the text line corresponding to the traversed optimized feature vector according to the nearest neighbor forward optimized feature vector and the nearest neighbor backward optimized feature vector.
[0135] In an alternative embodiment, determining the title level of the text line corresponding to the traversed optimized feature vector according to the nearest neighbor forward optimized feature vector and the nearest neighbor backward optimized feature vector includes:
[0136] In the case where there is no nearest neighbor forward optimized feature vector but there is a nearest neighbor backward optimized feature vector, determining whether the traversed optimized feature vector is the first ranked optimized feature vector in the third sorting result;
[0137] In the case where the traversed optimized feature vector is the optimized feature vector ranked first in the third sorting result, determine whether the first ratio between the font size feature in the traversed optimized feature vector and the font size feature in the nearest neighbor backward optimized feature vector is greater than a preset second font size threshold; in the case where the first ratio is greater than the preset second font size threshold, determine that the title level of the text line corresponding to the traversed optimized feature vector is a first-level title, and update the title level of the text line corresponding to the optimized feature vector with the serial number feature being the serial number value; or, in the case where the first ratio is not greater than the preset second font size threshold, determine the title level of the text line corresponding to the traversed optimized feature vector as the title level corresponding to the nearest neighbor backward optimized feature vector;
[0138] Or,
[0139] In the case where the traversed optimized feature vector is not the optimized feature vector ranked first in the third sorting result, find the previous optimized feature vector of the traversed optimized feature vector in the third sorting result; determine whether the second ratio between the font size feature in the previous optimized feature vector and the font size feature in the traversed optimized feature vector is greater than a preset third font size threshold; in the case where the second ratio is greater than the preset third font size threshold, determine the title level of the text line corresponding to the traversed optimized feature vector as the sum of the title level corresponding to the previous optimized feature vector and a preset level threshold; or, in the case where the second ratio is not greater than the preset third font size threshold, determine the title level of the text line corresponding to the traversed optimized feature vector as the title level corresponding to the previous optimized feature vector.
[0140] In an optional implementation manner, the method further includes:
[0141] In the case where there is a nearest neighbor forward optimized feature vector but no nearest neighbor backward optimized feature vector, in the third sorting result, determine whether the nearest neighbor forward optimized feature vector is the previous optimized feature vector of the position index;
[0142] In the case where the nearest neighbor forward optimized feature vector is the previous optimized feature vector of the position index, determine whether the third ratio between the font size feature in the traversed optimized feature vector and the font size feature in the nearest neighbor forward optimized feature vector is greater than a preset fourth font size threshold; in the case where the third ratio is greater than the preset fourth font size threshold, according to the document boundary feature in the traversed optimized feature vector, determine whether the text line corresponding to the traversed optimized feature vector is centered; in the case where the text line corresponding to the traversed optimized feature vector is centered, determine that the title level of the text line corresponding to the traversed optimized feature vector is a first-level title, and update the title level of the text line corresponding to the optimized feature vector with the serial number feature being the serial number value;
[0143] Or
[0144] When the nearest neighbor forward optimized feature vector is the previous optimized feature vector of the position index, determine whether the fourth ratio between the font size feature in the traversed optimized feature vector and the font size feature in the nearest neighbor forward optimized feature vector is less than the preset fifth font size threshold; when the fourth ratio is less than the preset fifth font size threshold, determine the title level of the text line corresponding to the traversed optimized feature vector as the sum of the title level corresponding to the nearest neighbor forward optimized feature vector and the preset level threshold; or, when the fourth ratio between the font size feature in the traversed optimized feature vector and the font size feature in the nearest neighbor forward optimized feature vector is not less than the preset fifth font size threshold, determine the title level of the text line corresponding to the traversed optimized feature vector as the title level corresponding to the nearest neighbor forward optimized feature vector;
[0145] Or
[0146] When the nearest neighbor forward optimized feature vector is not the previous optimized feature vector of the position index, determine whether the fifth ratio between the font size feature in the traversed optimized feature vector and the font size feature in the previous optimized feature vector is less than the preset sixth font size threshold; when the fifth ratio is less than the preset sixth font size threshold, determine the title level of the text line corresponding to the traversed optimized feature vector as the sum of the title level corresponding to the forward optimized feature vector and the preset level threshold; or, when the fifth ratio is not less than the preset sixth font size threshold, determine the title level of the text line corresponding to the traversed optimized feature vector as the title level corresponding to the forward optimized feature vector;
[0147] Or
[0148] In the case where there is a nearest neighbor forward optimization feature vector and a nearest neighbor backward optimization feature vector, determine whether the reference feature vector corresponding to the nearest neighbor forward optimization feature vector is the same as the reference feature vector corresponding to the nearest neighbor backward optimization feature vector; in the case of being the same, determine the sixth ratio between the font size feature in the traversed optimization feature vector and the font size feature in the nearest neighbor forward optimization feature vector; determine the magnitude between the serial number feature in the nearest neighbor forward optimization feature vector and the serial number feature in the nearest neighbor backward optimization feature vector; in the case where the sixth ratio is greater than the preset seventh font size threshold and the serial number feature in the nearest neighbor forward optimization feature vector is less than the serial number feature in the nearest neighbor backward optimization feature vector, determine the title level of the text line corresponding to the traversed optimization feature vector as the first-level title, and update the title level of the text line corresponding to the optimization feature vector with the serial number feature being the serial number value; or, in the case where the sixth ratio is not greater than the preset seventh font size threshold, and / or the serial number feature in the nearest neighbor forward optimization feature vector is not less than the serial number feature in the nearest neighbor backward optimization feature vector, determine the title level of the text line corresponding to the traversed optimization feature vector as the sum of the title level corresponding to the nearest neighbor forward optimization feature vector and the preset level threshold.
[0149] In an alternative embodiment, the method further includes:
[0150] In different cases, determine whether the title level corresponding to the nearest neighbor forward optimization feature vector is greater than the title level corresponding to the nearest neighbor backward optimization feature vector; in the case where the title level corresponding to the nearest neighbor forward optimization feature vector is greater than the title level corresponding to the nearest neighbor backward optimization feature vector, in the third sorting result, start from the nearest neighbor backward optimization feature vector and search backward for a new nearest neighbor backward optimization feature vector, and jump to the processing steps of the same case, where the reference feature vector of the new nearest neighbor backward optimization feature vector is the same as the reference feature vector corresponding to the nearest neighbor forward optimization feature vector;
[0151] Or, in the case where the title level corresponding to the nearest neighbor forward optimization feature vector is less than the title level corresponding to the nearest neighbor backward optimization feature vector, in the third sorting result, start from the nearest neighbor forward optimization feature vector and search forward for a new nearest neighbor forward optimization feature vector, and jump to the processing steps of the same case, where the reference feature vector of the new nearest neighbor forward optimization feature vector is the same as the reference feature vector corresponding to the nearest neighbor backward optimization feature vector;
[0152] Or, in the case where the title level corresponding to the nearest neighbor forward optimization feature vector is equal to the title level corresponding to the nearest neighbor backward optimization feature vector, determine the title level of the text line corresponding to the traversed optimization feature vector as the sum of the title level corresponding to the nearest neighbor forward optimization feature vector and the preset level threshold.
[0153] In the second aspect of the embodiments of the present application, a device for determining the document title hierarchy is further provided. The device includes:
[0154] A text line acquisition module, configured to acquire a target document, acquire text lines in a target document page of the target document, and store them in a title candidate line set;
[0155] A vector extraction module, configured to extract a basic feature vector and an optimized feature vector of each text line in the title candidate line set;
[0156] A title determination module, configured to determine a document title in the title candidate line set according to the basic feature vector and the optimized feature vector;
[0157] A title hierarchy determination module, configured to determine the title hierarchy of the document title.
[0158] In the third aspect of the embodiments of the present application, an electronic device is further provided, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus;
[0159] The memory is used to store a computer program;
[0160] The processor, when executing the program stored on the memory, implements the method for determining the document title hierarchy according to any one of the above first aspects.
[0161] In the fourth aspect of the embodiments of the present application, a storage medium is further provided. Instructions are stored in the storage medium, and when they run on a computer, the computer is made to execute the method for determining the document title hierarchy according to any one of the above first aspects.
[0162] In the fifth aspect of the embodiments of the present application, a computer program product containing instructions is further provided. When it runs on a computer, the computer is made to execute the method for determining the document title hierarchy according to any one of the above.
[0163] The technical solution provided by the embodiments of the present application acquires a target document, acquires text lines in a target document page of the target document, stores them in a title candidate line set, extracts a basic feature vector and an optimized feature vector of each text line in the title candidate line set, determines a document title in the title candidate line set according to the basic feature vector and the optimized feature vector, and determines the title hierarchy of the document title.
[0164] By extracting the basic feature vectors of each text line in the set of title candidate lines and optimizing the feature vectors, and based on the basic feature vectors and the optimized feature vectors, the document title in the set of title candidate lines is determined, and the title level of the document title is determined. Compared with the rule-based method, the recognition accuracy of the document title can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0165] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0166] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0167] One or more embodiments are exemplarily illustrated by the pictures in the corresponding accompanying drawings. These exemplary illustrations do not limit the embodiments. Elements with the same reference numerals in the accompanying drawings are represented as similar elements. Unless otherwise stated, the drawings in the accompanying drawings do not constitute a scale limitation.
[0168] Figure 1 It is a schematic flowchart of the implementation process of a method for determining the title level of a document title shown in the embodiments of the present application;
[0169] Figure 2 It is a schematic flowchart of the implementation process of another method for determining the title level of a document title shown in the embodiments of the present application;
[0170] Figure 3 It is a schematic flowchart of the implementation process of a method for training a bold classification model shown in the embodiments of the present application;
[0171] Figure 4 It is a schematic diagram of three relationships between the first character image and the second character image in a character image pair shown in the embodiments of the present application;
[0172] Figure 5 It is a schematic flowchart of the implementation process of a method for training a title classifier shown in the embodiments of the present application;
[0173] Figure 6 It is a schematic flowchart of the implementation process of a method for determining the title level of a document title with numbered characters shown in the embodiments of the present application;
[0174] Figure 7 It is a schematic flowchart of the implementation process of a method for determining the title level of a document title without numbered characters shown in the embodiments of the present application. Detailed implementation manners
[0175] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the scope of protection of the present application.
[0176] The following disclosure provides many different embodiments or examples for implementing different structures of the present application. To simplify the disclosure of the present application, components and settings of specific examples are described below. Of course, they are only examples and are not intended to limit the present application. In addition, the present application may repeat reference numerals and / or letters in different examples. This repetition is for the purpose of simplification and clarity and does not itself indicate the relationship between the various embodiments and / or settings discussed.
[0177] As Figure 1 shown, it is a schematic flowchart of an implementation process of a method for determining the level of a document title provided by an embodiment of the present application. This method is applied to an electronic device and may specifically include the following steps:
[0178] S101, obtain a target document, and obtain the text lines in the target document page of the target document, and store them in a title candidate line set.
[0179] In the embodiment of the present application, to obtain the target document, the target document may be a flowing document such as a Word document, or may be a layout document such as a PDF document or an OFD document, or may be an HTML document. The embodiment of the present application does not limit this.
[0180] In an embodiment of the present application, for the target document, obtain the text lines in the target document page of the target document, and store them in a title candidate line set. Among them, for the target document page, it may refer to a specific page or several document pages in the target document. Of course, it may also be all document pages. In this way, a specific page or several document pages of the target document, or all document pages can be obtained.
[0181] In another embodiment of the present application, the text lines in the target document page of the target document may also be the text lines in a specified interval or meeting a predetermined condition in the target document page. For example, the text line at the beginning of each paragraph in the target document page and the next adjacent text line.
[0182] In addition, for text lines, it can also refer to all text lines in the target document page that are not in tables, headers, or footers, thereby obtaining a set of title candidate lines {page1(line1~n), page2(line1~n), ……, pagen(line1~n)}, where page1(line1~n) represents the first page of the document page, which contains text lines {line1, line2, ……, linen}.
[0183] S102. Extract the basic feature vectors and optimized feature vectors of each text line in the set of title candidate lines.
[0184] In an embodiment of the present application, for the set of title candidate lines obtained in the above steps, extract the basic feature vectors and optimized feature vectors of each text line in the set of title candidate lines. Among them, the basic feature vector is represented by LineFeature, and the optimized feature vector is represented by Feature.
[0185] S103. Determine the document title in the set of title candidate lines according to the basic feature vectors and optimized feature vectors, and determine the title level of the document title.
[0186] In an embodiment of the present application, through the above steps, the basic feature vectors and optimized feature vectors of each text line in the set of title candidate lines can be extracted. Thus, according to the basic feature vectors and optimized feature vectors, determine the document title in the set of title candidate lines, and determine the title level of the document title.
[0187] In this way, extract the basic feature vectors and optimized feature vectors of each text line in the set of title candidate lines, and determine the document title in the set of title candidate lines according to the basic feature vectors and optimized feature vectors. Compared with the rule-based method, the recognition accuracy of the document title can be improved.
[0188] In addition, in an embodiment of the present application, for the optimized feature vector of a text line, it is composed of optimized features, and the optimized features include at least one of the following: font application feature, bold feature, serial number feature, font size feature, document boundary feature, page index feature, line index feature, text feature, color feature. For the extraction of the basic feature vector of a text line, reference can be made to the prior art.
[0189] Based on this, as Figure 2 shown, it is a schematic flowchart of the implementation process of another method for determining the document title level provided by an embodiment of the present application. This method is applied to an electronic device and specifically may include the following steps:
[0190] S201. Obtain the target document, and obtain the text lines in the target document page of the target document, and store them in the set of title candidate lines.
[0191] In the embodiment of the present application, this step is similar to the above step S101, and the embodiments of the present application will not be elaborated herein one by one.
[0192] S202. Extract the basic feature vectors and optimized features of each text line in the title candidate line set, and form an optimized feature vector from the optimized features.
[0193] In the embodiment of the present application, the basic feature vectors of each text line in the title candidate line set are extracted. In addition, the optimized features of each text line in the title candidate line set are also extracted, and an optimized feature vector is formed from the optimized features. The optimized features include at least one of the following: font application feature, bold feature, serial number feature, font size feature, document boundary feature, page index feature, line index feature, text feature, color feature.
[0194] Among them, for the optimized feature, when the optimized feature is the font application feature, the font application feature is extracted according to the font application situation of the characters in the text line. If all the characters in the current text line are written and drawn using the same font, then the font application feature feature【0】 can be set to, for example, 1, otherwise it is set to, for example, 0, indicating the existence of mixed fonts.
[0195] Based on this, for any text line in the title candidate line set, obtain the fonts corresponding to each character in the text line. When all the fonts are the same, set the font application feature to the first value (for example, the first value is 1). When all or some of the fonts are different, set the font application feature to the second value (for example, the second value is 0).
[0196] It should be noted that for obtaining the fonts corresponding to each character in the text line, in the case where the target document is a flow document, directly obtain the fonts corresponding to each character in the text line. In the case where the target document is a layout document, obtain the fonts recorded in the text objects corresponding to each character in the text line. The embodiments of the present application do not make any limitations in this regard.
[0197] For the optimized feature, when the optimized feature is the bold feature feature【1】, the bold feature can be extracted according to the font usage situation or font weight of the characters in the current text line. Additionally, in the case where the above conditions cannot be used for extraction, when the current text line is a non-mixed font, it is necessary to render the current text line and the next text line as images and extract the bold feature from the images. When the current text line is a mixed font, it is necessary to render by font, render the content of the same font onto the same image, and then extract the bold feature from the image rendered with the next text line.
[0198] Based on this, for any text line in the set of title candidate lines, obtain the fonts corresponding to each character in the text line. When all the fonts are the same and are bold fonts, set the bold feature to a first value (for example, the first value is 1). Or, for any text line in the set of title candidate lines, obtain the fonts and font weights corresponding to each character in the text line. When all the fonts are the same and all the font weights are greater than a preset weight threshold (for example, the preset weight threshold is 700), set the bold feature to the first value.
[0199] In addition, when all the fonts are the same, all the font weights are not greater than the preset weight threshold, and all the fonts are not bold fonts, perform the following processing: Render the text line to obtain a first text line image, and find the next text line of the text line from the set of title candidate lines; Render the next text line to obtain a second text line image, and set the bold feature according to the first text line image and the second text line image.
[0200] In addition, when all or some of the fonts are different, all the font weights are not greater than the preset weight threshold, and all the fonts are not bold fonts, perform the following processing: Classify the text included in the text line according to the font types included in the text line, and set the bold feature for the text included in the text line for different font types.
[0201] Among them, classifying the text included in the text line according to the font types included in the text line and setting the bold feature for the text included in the text line for different font types specifically means rendering the text line by font to obtain at least two third text line images, and finding the next text line of the text line from the set of title candidate lines; Render the next text line to obtain a second text line image, and set the bold feature according to the at least two third text line images and the second text line image.
[0202] In the embodiments of the present application, a pre-trained bold classification model is introduced. The above first text line image and second text line image, or at least two third text line images and second text line image can be input into the pre-trained bold classification model to obtain a model output result, and the bold feature is set according to the model output result. The model output result can be divided into 3 categories, for example, {0: reverse bold, 1: non-bold, 2: bold}. Among them, for reverse bold, taking the current text line and the next text line as an example, it means that the current text line is not bold, while the next text line is bold, and a reverse bold is formed between the two (that is, the current text line is thinner than the next text line).
[0203] Based on this, the first text line image and the second text line image are input into a pre-trained bold classification model to obtain a first output result. When the first output result is a preset result (for example, the preset result is 2), the bold feature is set to a first value; or when the first output result is not the preset result, the bold feature is set to a second value.
[0204] In addition, for any third text line image, the third text line image and the second text line image are input into a pre-trained bold classification model to obtain a second output result; when all the second output results are preset results, the bold feature is set to a first value; or when all or some of the second output results are not preset results, the bold feature is set to a second value.
[0205] For example, for the text line "Chapter 1 xxx", the font of "Chapter 1" is boldface, and the font of "xxx" is Song typeface. All font weights are not greater than a preset weight threshold and all fonts are not bold fonts. Then, it is necessary to render the fonts separately to obtain two third text line images, that is, the first third text line image corresponds to "Chapter 1", and the second third text line image corresponds to "xxx". Then, they are respectively used as inputs to a pre-trained bold classification model together with the second text line image to obtain an output set (r1, r2). When all the output set is 2, the bold feature feature【1】 is set to 1, otherwise it is set to 0.
[0206] In the embodiments of the present application, for the pre-trained bold classification model, it needs to be trained to obtain. Based on this, as Figure 3 shown, it is a schematic flowchart of an implementation process of a training method for a bold classification model provided by an embodiment of the present application. This method is applied to an electronic device and specifically may include the following steps:
[0207] S301, obtain a character image pair and a sample label corresponding to the character image pair, input the character image pair into the bold classification model, and obtain a prediction result.
[0208] In the embodiments of the present application, a number of non-commercial fonts are collected, and the characters in the fonts are rendered in the form of images, and then a character image pair is constructed. Among them, the two characters in the character image pair use different fonts, and there are mainly the following three relationships in the character image pair: the first character image is thicker than the second character image, the first character image is as thick as the second character image, and the first character image is thinner than the second character image.
[0209] For convenience, the above three relationships are simplified, and bold, non-bold, and reverse bold are used to represent these three relationships. For these three relationships, as Figure 4As shown. And for these three relationships, corresponding sample labels can be set. For example, reverse bold corresponds to sample label 0, non-bold corresponds to sample label 1, and bold corresponds to sample label 2.
[0210] In this way, character image pairs and the sample labels corresponding to the character image pairs can be obtained. Among them, for a character image pair including a first character image and a second character image, the character image pair, that is, the first character image and the second character image, can be input into the bold classification model to obtain a prediction result.
[0211] It should be noted that for the bold classification model, a machine learning classification model such as SVM can be selected, or a classification model of deep learning can also be used. The embodiments of the present application do not limit this.
[0212] S302, determine the loss value between the prediction result and the sample label, and train the bold classification model according to the loss value.
[0213] S303, when the loss value converges, stop training to obtain a pre-trained bold classification model.
[0214] In the embodiments of the present application, determine the loss value between the prediction result and the sample label, and train the bold classification model according to the loss value. When the loss value converges (for example, the loss value is less than a certain threshold), stop training to obtain a pre-trained bold classification model.
[0215] It should be noted that in order to increase the diversity of training data and improve the accuracy of final recognition, the character image pairs can be exchanged, that is, the second character image is used as the first character image, the first character image is used as the second character image, and the corresponding sample labels are transformed, and then input into the bold classification model for training.
[0216] For example, if a character image pair includes a first character image and a second character image, and its corresponding sample label is 2, then the second character image is used as the first character image, the first character image is used as the second character image, and the corresponding sample label is transformed into 0, and then input into the bold classification model for training.
[0217] In addition, for the optimized feature, when the optimized feature is the serial number feature feature【2】, the serial number feature is extracted according to the serial number character at the beginning of the current text line. Based on this, for any text line in the title candidate line set, detect whether there is a serial number character at the beginning of the text line. If there is a serial number character, find the serial number value corresponding to the serial number character, and set the serial number feature to the serial number value. Or, if there is no serial number character, set the serial number feature to the second value.
[0218] For example, for any text line in the set of title candidate lines, detect whether there is a serial number character at the beginning of the text line. For serial number characters, they can be, for example, "1.", "2.1", "one", "Chapter 1", etc. In the case where there is a serial number character, find the serial number value corresponding to the serial number character. For example, "1." represents the first category, so its serial number value is 1, and "(1)" represents the second category, so its serial number value is 2. Thus, set the serial number feature to the serial number value. In the case where there is no serial number character, set the serial number feature to 0.
[0219] In addition, for the optimization feature, in the case where the optimization feature is the font size feature feature【3】, calculate the average font size of the current text line and set it as the font size feature feature【3】. Based on this, for any text line in the set of title candidate lines, obtain the font sizes corresponding to each character in the text line; determine the average font size of the font sizes and set the font size feature to the average font size.
[0220] It should be noted that in the process of calculating the average font size, the basic punctuation symbol set needs to be excluded, that is, basic punctuation symbols (such as commas, periods) do not participate in the calculation of the average font size. And in the case where the target document is a flowing document, directly obtain the font sizes corresponding to each character in the text line. In the case where the target document is a layout document, query the font sizes recorded in the text objects corresponding to each character in the text line. In the case where the target document is neither a flowing document nor a layout document, the font size of the character can be obtained by extracting the text box size of the character. The embodiments of the present application do not make any limitations in this regard.
[0221] In addition, for the optimization feature, in the case where the optimization feature is the document boundary feature, calculate the indentation of the text line relative to the document boundary (that is, the distance of the text line relative to the document boundary) and set it as the document boundary feature feature[4]. Based on this, for any text line in the set of title candidate lines, determine the distance between the text line and the document boundary and set the document boundary feature to the distance.
[0222] It should be noted that for the document boundary, it is usually divided into the left document boundary and the right document boundary. Then the above distance can be further divided into the left distance between the text line and the left document boundary and the right distance between the text line and the right document boundary. The embodiments of the present application do not make any limitations in this regard.
[0223] In addition, for the optimized features, when the optimized features are page index features and line index features, the page index feature feature[5] and the line index feature feature[6] are set respectively according to the page index and the line index. Among them, for the page index, it can be understood as which page of the target document the current text line is located on, and for the line index, it can be understood as which line of the current document page in the target document the current text line is located on.
[0224] Based on this, for any text line in the title candidate line set, determine the page index and line index of the text line, set the page index feature to the page index, and set the line index feature to the line index. In this way, for each text line, it can be known which page of the target document it is located on and which line it is on.
[0225] In addition, for the optimized features, when the optimized feature is a text feature, set the characters included in the current text line as the text feature feature[7]. Based on this, for any text line in the title candidate line set, obtain the characters included in the text line, and set the text feature to the characters included in the text line.
[0226] In addition, for the optimized features, when the optimized feature is a color feature, the current text line can be rendered as an image, and the character region in the image is extracted. Among them, in one embodiment of the present application, after image processing such as binarization operation, morphological operation, and edge extraction, the character region can be extracted. Calculate the pixel average value of the character region and set it as the color feature feature[8]. Based on this, for any text line in the title candidate line set, determine the character region in the text line, determine the pixel average value of the character region, and set the color feature to the pixel average value.
[0227] It should be noted that for the character region, it can be a grayscale image, then the pixel average value is a single value, or it can be a color image, then the pixel average values respectively correspond to the three RGB channels, and their value ranges are in [0, 255]. The embodiments of the present application do not limit this.
[0228] S203. For any text line in the title candidate line set, input the basic feature vector and the optimized feature vector of the text line into the pre-trained title classifier to obtain a classification probability value.
[0229] In the embodiment of the present application, for any text line in the title candidate line set, input the basic feature vector and the optimized feature vector of the text line into the pre-trained title classifier to obtain a classification probability value.
[0230] Among them, for the pre-trained title classifier, it needs to be trained to obtain. Based on this, as Figure 5As shown in the figure, it is a schematic flowchart of the implementation process of a training method for a title classifier provided by an embodiment of the present application. This method is applied to an electronic device and may specifically include the following steps:
[0231] S501, Obtain a sample title corpus, and extract the sample basic feature vector and the sample optimized feature vector of the sample title corpus.
[0232] In the embodiment of the present application, a sample title corpus is obtained, and the sample basic feature vector and the sample optimized feature vector of the sample title corpus are extracted. Among them, for the extraction of the sample optimized feature vector, it is similar to the extraction of the above-mentioned optimized feature vector, and the embodiments of the present application will not elaborate here one by one.
[0233] It should be noted that the sample title corpus can be replaced with a sample text corpus, and correspondingly, the sample basic feature vector and the sample optimized feature vector of the sample text corpus are extracted. The embodiments of the present application do not limit this.
[0234] S502, Input the sample basic feature vector and the sample optimized feature vector into the title classifier to obtain a predicted classification probability value.
[0235] In the embodiment of the present application, for the extracted sample basic feature vector and sample optimized feature vector, the sample basic feature vector and the sample optimized feature vector can be input into the title classifier to obtain a predicted classification probability value.
[0236] It should be noted that the title classifier includes, but is not limited to, logistic regression, support vector machine, neural network, decision tree, random forest, Bayesian classifier, etc. The embodiments of the present application do not limit this.
[0237] S503, Determine the probability loss between the predicted classification probability value and the sample classification probability value corresponding to the sample title corpus.
[0238] S504, Train the title classifier according to the probability loss, and stop training when the probability loss converges to obtain a pre-trained title classifier.
[0239] In the embodiment of the present application, determine the probability loss between the predicted classification probability value and the sample classification probability value (for example, the sample classification probability value is 1 or 0) corresponding to the sample title corpus (or sample text corpus).
[0240] Train the title classifier according to the probability loss, and stop training when the probability loss converges (for example, the probability loss is less than a certain threshold) to obtain a pre-trained title classifier.
[0241] S204. When the classification probability value is greater than the preset classification threshold, determine the text line as the document title, and store the optimized feature vector of the text line into the title feature set.
[0242] In the embodiment of the present application, for the classification probability value obtained in the above step, when the classification probability value is greater than the preset classification threshold (for example, the preset classification threshold is 0.75), determine the text line as the document title, and store the optimized feature vector of the text line into the title feature set.
[0243] For example, when the classification probability value conf is greater than 0.75, then determine the text line as the document title, and store the optimized feature vector of the text line into the title feature set vecHeadings, which means that the title feature set vecHeadings can be regarded as storing the optimized feature vectors of document titles.
[0244] S205. When the classification probability value is not greater than the preset classification threshold, obtain the serial number feature or the font size feature in the optimized feature vector of the text line.
[0245] S206. When the serial number feature is not the second value or the font size feature is greater than the preset first font size threshold, store the optimized feature vector of the text line into the candidate title feature set.
[0246] In the embodiment of the present application, for the classification probability value obtained in the above step, when the classification probability value is not greater than the preset classification threshold, at this time, for the text line, it may be a suspected document title and needs further confirmation. For this reason, store the optimized feature vector of the text line that meets the first preset condition into the candidate title feature set, and determine the title level of the document title according to the title feature set and the candidate title feature set.
[0247] Among them, the above first preset condition refers to that the serial number feature is not the second value (for example, the second value is 0) or the font size feature is greater than the preset first font size threshold (for example, the preset first font size threshold is 18). For this reason, obtain the serial number feature or the font size feature in the optimized feature vector of the text line. When the serial number feature is not the second value (for example, the second value is 0) or the font size feature is greater than the preset first font size threshold (for example, the preset first font size threshold is 18), determine the text line as a suspected document title. At this time, store the optimized feature vector of the text line into the candidate title feature set vecDistrustHeadings for subsequent confirmation, so as to supplement the document titles missed in the previous step as much as possible.
[0248] S207. Divide the title feature set into a first title feature set and a second title feature set.
[0249] In an embodiment of the present application, for the title feature set vecHeadings obtained in the above steps, the title feature set vecHeadings can be divided into a first title feature set vecSerialHeadings and a second title feature set vecNoSerialHeadings.
[0250] Among them, for the division rule of the title feature set vecHeadings, it can be divided into a first title feature set vecSerialHeadings and a second title feature set vecNoSerialHeadings according to the serial number feature in the optimized feature vector.
[0251] Specifically, traverse the optimized feature vector of the title feature set to obtain the serial number feature in the traversed optimized feature vector; when the serial number feature is not the second value, store the optimized feature vector in the first title feature set; when the serial number feature is the second value, store the optimized feature vector in the second title feature set.
[0252] It should be noted that for the first title feature set vecSerialHeadings, it represents the text lines confirmed as document titles and having serial number characters, and for the second title feature set vecNoSerialHeadings, it represents the text lines confirmed as document titles but without serial number characters.
[0253] S208. Divide the candidate title feature set into a first candidate title feature set and a second candidate title feature set.
[0254] In an embodiment of the present application, for the candidate title feature set vecDistrustHeadings obtained in the above steps, the candidate title feature set vecDistrustHeadings can be divided into a first candidate title feature set vecSerialDistrustHeadings and a second candidate title feature set vecNoSerialDistrustHeadings.
[0255] Among them, for the division rule of the candidate title feature set vecDistrustHeadings, it can be divided into a first candidate title feature set vecSerialDistrustHeadings and a second candidate title feature set vecNoSerialDistrustHeadings according to the serial number feature in the optimized feature vector.
[0256] Specifically, traverse the optimized feature vectors in the candidate title feature set, and obtain the serial number features in the traversed optimized feature vectors; when the serial number feature is not the second value, store the optimized feature vector in the first candidate title feature set; when the serial number feature is the second value, store the optimized feature vector in the second candidate title feature set.
[0257] It should be noted that for the first candidate title feature set vecSerialDistrustHeadings, it represents the text lines suspected of being document titles and there are serial number characters, and for the second candidate title feature set vecNoSerialDistrustHeadings, it represents the text lines suspected of being document titles but there are no serial number characters.
[0258] S209. Determine the first candidate optimized feature vector in the first candidate title feature set, and store it in the first title feature set to obtain the first target title feature set, and determine the title level of the text line corresponding to the optimized feature vector in the first target title feature set.
[0259] In the embodiments of the present application, the title level of the document title with serial number characters is calculated preferentially. Thus, determine the first candidate optimized feature vector in the first candidate title feature set, and store it in the first title feature set to obtain the first target title feature set, and determine the title level of the text line corresponding to the optimized feature vector in the first target title feature set.
[0260] Among them, cluster the optimized feature vectors in the first title feature set to obtain several clustering sets, determine the benchmark feature vector corresponding to each clustering set, and then traverse the optimized feature vectors in the first candidate title feature set. If it matches a certain benchmark feature vector, store it in the clustering set corresponding to the benchmark feature vector. In this way, several clustering sets are updated, and the updated several clustering sets are merged to obtain the first target title feature set, and determine the title level of the text line corresponding to the optimized feature vector in the updated several clustering sets.
[0261] Based on this, for the above step S209, the method shown in Figure 6 can be referred to. As shown in Figure 6 is the implementation process schematic diagram of a method for determining the title level of a document title with serial number characters provided by the embodiments of the present application. This method is applied to an electronic device and specifically may include the following steps:
[0262] S601. Cluster the optimized feature vectors in the first title feature set to obtain multiple first clustering sets, and determine the benchmark feature vector corresponding to each first clustering set.
[0263] In the embodiment of the present application, for the first heading feature set vecSerialHeadings obtained in the above steps, the optimized feature vectors in the first heading feature set vecSerialHeadings can be clustered to obtain multiple first clustering sets, and the benchmark feature vector corresponding to each first clustering set is determined. Among them, the benchmark feature vector is similar to the above optimized feature vector.
[0264] Among them, for the optimized feature vectors in the first heading feature set vecSerialHeadings, clustering is preferably performed according to the serial number features. Secondly, if the serial number features are the same, clustering is performed according to the font size features. During the clustering process according to the font size features, if the difference between the font size feature in the current optimized feature vector and the average value of the font size features of the optimized feature vectors in a certain clustering set is too large (for example, the gap is 1.5 times), then the current optimized feature vector should belong to a new clustering set.
[0265] Based on this, obtain the serial number features included in the optimized feature vectors in the first heading feature set, and determine whether the serial number features are the same. In the case where the serial number features are different, cluster the optimized feature vectors in the first heading feature set according to the serial number features to obtain multiple first clustering sets. In the case where the serial number features are the same, cluster the optimized feature vectors in the first heading feature set according to the font size features to obtain multiple first clustering sets.
[0266] For example, classify "Chapter x", "Section x", and "1.xxx" into three different categories, and their corresponding serial number values are 1, 2, and 3 respectively. Then, in the case where the serial number features are different, cluster the optimized feature vectors in the first heading feature set according to the serial number features to obtain 3 first clustering sets. The first clustering set 1 corresponds to the optimized feature vectors with serial number feature 1, the first clustering set 2 corresponds to the optimized feature vectors with serial number feature 2, and the first clustering set 3 corresponds to the optimized feature vectors with serial number feature 3.
[0267] In addition, in the embodiment of the present application, the benchmark feature vector is similar to the above optimized feature vector and can represent the common attributes of the optimized feature vectors in the corresponding first clustering set. Thus, for any first clustering set, extract the benchmark features corresponding to the first clustering set, and form a benchmark feature vector from the benchmark features; among them, the benchmark features include at least one of the following: benchmark font application feature, benchmark bold feature, benchmark serial number feature, benchmark font size feature, benchmark document boundary feature, benchmark page index feature, benchmark line index feature, benchmark text feature, benchmark color feature.
[0268] It should be noted that for the benchmark feature vector, a corresponding pair can be formed with the corresponding first clustering set, and the following structure benchmarki can be obtained: the first clustering set i, that is, the i-th clustering set and the i-th benchmark form a corresponding pair, which is convenient for subsequent queries. The embodiments of the present application do not limit this.
[0269] Among them, for the benchmark feature, when the benchmark feature is the benchmark font application feature, count the first font application quantity corresponding to the optimized feature vector with the font application feature being the first value in the first clustering set; count the second font application quantity corresponding to the optimized feature vector with the font application feature being the second value in the first clustering set; when the first font application quantity is greater than the second font application quantity, set the benchmark font application feature corresponding to the first clustering set to the first value; or, when the first font application quantity is not greater than the second font application quantity, set the benchmark font application feature corresponding to the first clustering set to the second value.
[0270] For the benchmark feature, when the benchmark feature is the benchmark bold feature, count the first bold quantity corresponding to the optimized feature vector with the bold feature being the first value in the first clustering set; count the second non-bold quantity corresponding to the optimized feature vector with the bold feature being the second value in the first clustering set; when the first bold quantity is greater than the second non-bold quantity, set the benchmark bold feature corresponding to the first clustering set to the first value; or, when the first bold quantity is not greater than the second non-bold quantity, set the benchmark bold feature corresponding to the first clustering set to the second value.
[0271] For the benchmark feature, when the benchmark feature is the benchmark serial number feature, obtain the serial number feature included in the optimized feature vector in the first clustering set, and set the benchmark serial number feature corresponding to the first clustering set to the serial number feature.
[0272] For the benchmark feature, when the benchmark feature is the benchmark font size feature, obtain the font size feature included in the optimized feature vector in the first clustering set, and determine the average font size feature of the font size feature; set the benchmark font size feature corresponding to the first clustering set to the average font size feature.
[0273] For the benchmark feature, when the benchmark feature is the benchmark document boundary feature, obtain the document boundary feature included in the optimized feature vector in the first clustering set, and determine the average document boundary feature of the document boundary feature; set the benchmark document boundary feature corresponding to the first clustering set to the average document boundary feature.
[0274] It should be noted that the document boundary features can generally be subdivided into the left document boundary feature and the right document boundary feature, which respectively refer to the distance between the text line and the left document boundary, and the distance between the text line and the right document boundary. From this, the average left document boundary feature of the document left boundary feature and the average right document boundary feature of the document right boundary feature are calculated. The reference document left boundary feature corresponding to the first clustering set is set as the average left document boundary feature, and the reference document right boundary feature corresponding to the first clustering set is set as the average right document boundary feature.
[0275] For the reference feature, when the reference feature is the reference page index feature or the reference line index feature, obtain the page index features included in the optimized feature vector in the first clustering set, and select the smallest page index feature from the page index features; set the reference page index feature corresponding to the first clustering set as the smallest page index feature; obtain the line index features included in the optimized feature vector in the first clustering set, and select the smallest line index feature from the line index features; set the reference line index feature corresponding to the first clustering set as the smallest line index feature.
[0276] For the reference feature, when the reference feature is the reference text feature, set the reference text feature corresponding to the first clustering set as empty, and empty means not filling in any content.
[0277] For the reference feature, when the reference feature is the reference color feature, obtain the color features included in the optimized feature vector in the first clustering set, and determine the average color feature of the color features; set the reference color feature corresponding to the first clustering set as the average color feature.
[0278] S602, traverse the optimized feature vectors in the first candidate title feature set, and determine the first similarity between the traversed optimized feature vector and any reference feature vector.
[0279] In the embodiment of the present application, for the first candidate title feature set vecSerialDistrustHeadings, traverse the optimized feature vectors in the first candidate title feature set vecSerialDistrustHeadings. For the traversed optimized feature vector, determine the first similarity between it and any of the above-determined reference feature vectors.
[0280] It should be noted that in the process of calculating the first similarity between the optimized feature vector and the reference feature vector, the difference between the same type of features can be calculated, and then weighted and summed. For example, the optimized feature vector is composed of the above 9 optimized features, and the corresponding reference feature vector is also composed of 9 reference features. The difference between the font application feature and the reference font application feature can be calculated, and the rest are similar, and then weighted and summed.
[0281] In addition, during the process of calculating the first similarity between the optimized feature vector and the reference feature vector, some features can also be involved in the calculation. For example, feature[0] to feature[3] in the optimized feature vector are involved in the calculation, and the corresponding reference features in the reference feature vector corresponding to feature[0] to feature[3] are involved in the calculation.
[0282] S603. Find the maximum first target similarity from the first similarities, and determine whether the first target similarity is greater than a preset first similarity threshold.
[0283] S604. In the case where the first target similarity is greater than the preset first similarity threshold, determine the traversed optimized feature vector as the first candidate optimized feature vector.
[0284] In the embodiments of the present application, through the above steps, multiple first similarities can be obtained. The maximum first target similarity can be found from these first similarities, and it can be determined whether the first target similarity is greater than the preset first similarity threshold. In the case where the first target similarity is greater than the preset first similarity threshold, the traversed optimized feature vector is determined as the first candidate optimized feature vector, indicating the document title with missed detection of the text behavior corresponding to the traversed optimized feature vector.
[0285] Among them, for the first candidate optimized feature vector, the first candidate optimized feature vector can be stored in the first title feature set to obtain the first target title feature set, and the title level of the text line corresponding to the optimized feature vector in the first target title feature set is determined. Specifically, after determining the traversed optimized feature vector as the first candidate optimized feature vector, determine the reference feature vector corresponding to the first target similarity, and find the first clustering set corresponding to the determined reference feature vector. Store the first candidate optimized feature vector in the found first clustering set, and determine the title level of the text line corresponding to the optimized feature vector in each updated first clustering set; among them, the updated first clustering sets are merged to obtain the first target title feature set, that is, refer to the subsequent steps S605 - S606.
[0286] S605. Determine the reference feature vector corresponding to the first target similarity, find the first clustering set corresponding to the determined reference feature vector, and store the first candidate optimized feature vector in the found first clustering set.
[0287] In the embodiments of the present application, for the first target similarity obtained in the above steps, the reference feature vector corresponding to the first target similarity can be determined, the first clustering set corresponding to the determined reference feature vector can be found, and the first candidate optimized feature vector can be stored in the found first clustering set, thus completing the supplement of the missed detection document title.
[0288] It should be noted that for the first clustering set, after the above supplementation, the first candidate optimization feature vector is added. For the updated first clustering set, the benchmark page index and benchmark line index in the corresponding benchmark feature vector also need to be updated, and the update method is similar to the setting of the above benchmark page index and benchmark line index.
[0289] S606. Determine the title level of the text line corresponding to the optimization feature vector in each updated first clustering set; wherein, each updated first clustering set is merged to obtain the first target title feature set.
[0290] In the embodiment of the present application, for each first clustering set, after the above-mentioned missed document title supplementation step, an update is obtained, and thus the title level of the text line corresponding to the optimization feature vector in each updated first clustering set is determined; wherein, each updated first clustering set is merged to obtain the first target title feature set vecAllSerialHeadings.
[0291] Among them, for the benchmark feature vector, the level of each benchmark feature vector can represent the title level of the text line corresponding to the optimization feature vector in the corresponding updated first clustering set. Therefore, for all benchmark feature vectors, a first sorting result is obtained through sorting. According to the first sorting result, the title level of the text line corresponding to the optimization feature vector in each updated first clustering set is determined.
[0292] It should be noted that when sorting the benchmark feature vectors, the sorting can be performed according to the benchmark page index feature and benchmark line index feature included in the benchmark feature vector. Among them, the sorting is preferably performed in ascending order according to the benchmark page index feature. During this process, if the benchmark page index features are the same, then the sorting is performed in ascending order according to the benchmark line index feature. The embodiment of the present application does not make any limitations in this regard.
[0293] Specifically, traverse the benchmark feature vectors in the first sorting result in sequence. For the traversed benchmark feature vector, in the case where it is the first benchmark feature vector (i.e., the benchmark feature vector ranked first in the first sorting result), search for the updated first clustering set corresponding to the traversed benchmark feature vector, and determine the first-level title corresponding to the traversed benchmark feature vector as the title level of the text line corresponding to the optimization feature vector in the updated first clustering set corresponding to the traversed benchmark feature vector.
[0294] In addition, when the traversed reference feature vector is not the first reference feature vector (i.e., the reference feature vector ranked first in the first sorting result), find the previous reference feature vector of the traversed reference feature vector (i.e., the previous reference feature vector before the currently traversed reference feature vector in the first sorting result) from the first sorting result, and determine the title level of the text line corresponding to the optimized feature vector in the updated first clustering set corresponding to the traversed reference feature vector according to the previous reference feature vector.
[0295] Among them, for the reference feature vectors traversed in sequence in the first sorting result, they need to be compared with the previous reference feature vector. If the reference page indexes included in both are the same but the reference line indexes are different, since the previous reference feature vector is in front of the traversed reference feature vector, the title level of the traversed reference feature vector is the title level of the previous reference feature vector + 1, and the traversed reference feature vector is added to the sub-title list of the previous reference feature vector.
[0296] Based on this, when the reference page index feature included in the previous reference feature vector is the same as the reference page index feature included in the traversed reference feature vector, determine the previous title level corresponding to the previous reference feature vector, and determine the sum of the previous title level and the preset level threshold (for example, the preset level threshold is 1) as the title level corresponding to the traversed reference feature vector, determine the title level corresponding to the traversed reference feature vector as the title level of the text line corresponding to the optimized feature vector in the updated first clustering set corresponding to the traversed reference feature vector, and add the traversed reference feature vector to the sub-title list of the previous reference feature vector.
[0297] In addition, for the reference feature vectors traversed in sequence in the first sorting result, compare them with the previous reference feature vector. If the reference page indexes included in both are different, then for the updated first clustering set corresponding to the traversed reference feature vector, traverse the optimized feature vectors in the updated first clustering set to determine the title level corresponding to the traversed reference feature vector.
[0298] Specifically, when the reference page index feature included in the previous reference feature vector is different from the reference page index feature included in the traversed reference feature vector, sort the optimized feature vectors in each updated first clustering set to obtain a second sorting result; for the updated first clustering set corresponding to the traversed reference feature vector, traverse the optimized feature vectors in the updated first clustering set, and from the second sorting result, find the previous optimized feature vector of the traversed optimized feature vector (that is, from the second sorting result, find the position index of the traversed optimized feature vector, and then determine the previous optimized feature vector as the previous optimized feature vector).
[0299] In the case where there is no previous optimized feature vector, from the second sorting result, find the next optimized feature vector of the traversed optimized feature vector (that is, from the second sorting result, find the position index of the traversed optimized feature vector, and then determine the next optimized feature vector as the next optimized feature vector). In the case where there is a next optimized feature vector, find the reference feature vector corresponding to the updated first clustering set where the next optimized feature vector is located.
[0300] When the traversed reference feature vector is different from the reference feature vector corresponding to the updated first clustering set where the next optimized feature vector is located, and the product of the reference font size feature included in the next optimized feature vector and a preset first multiple (for example, the preset first multiple is 1.2) is less than the reference font size feature included in the traversed reference feature vector, determine the next heading level of the reference feature vector corresponding to the updated first clustering set where the next optimized feature vector is located.
[0301] Determine the difference between the next heading level and a preset level threshold (for example, the preset level threshold is 1) as the heading level corresponding to the traversed reference feature vector; determine the heading level corresponding to the traversed reference feature vector as the heading level of the text line corresponding to the optimized feature vector in the updated first clustering set corresponding to the traversed reference feature vector, add the reference feature vector corresponding to the updated first clustering set where the next optimized feature vector is located to the sub-heading list of the traversed reference feature vector, and for the reference feature vectors whose heading levels have been determined, their corresponding heading levels need to be updated (for example, the next heading level of the reference feature vector corresponding to the updated first clustering set where the next optimized feature vector is located + 1, and correspondingly, the heading levels of the reference feature vectors in the sub-heading list of the reference feature vector corresponding to the updated first clustering set where the next optimized feature vector is located + 1).
[0302] In addition, in the presence of a previous optimized feature vector, determine the previous title level of the reference feature vector corresponding to the updated first clustering set where the previous optimized feature vector is located; determine the sum of the previous title level and a preset level threshold (for example, the preset level threshold is 1) as the title level corresponding to the traversed reference feature vector; determine the title level corresponding to the traversed reference feature vector as the title level of the text line corresponding to the optimized feature vector in the updated first clustering set corresponding to the traversed reference feature vector; add the traversed reference feature vector to the sub - title list of the subsequent reference feature vector.
[0303] After the above steps, for each reference feature vector, there is a corresponding title level. The title level of each reference feature vector can represent the title level of the text line corresponding to the optimized feature vector in the corresponding updated first clustering set. Thus, for the text lines (i.e., document titles) corresponding to the optimized feature vectors in each updated first clustering set, there are corresponding title levels.
[0304] S210. Determine the second candidate optimized feature vectors in the second candidate title feature set, store them in the second title feature set to obtain the second target title feature set, and determine the title level of the text line corresponding to the optimized feature vector in the second target title feature set.
[0305] In the embodiment of the present application, calculate the title level of the document title without serial number characters. Thus, determine the second candidate optimized feature vectors in the second candidate title feature set, store them in the second title feature set to obtain the second target title feature set, and determine the title level of the text line corresponding to the optimized feature vector in the second target title feature set.
[0306] Among them, cluster the optimized feature vectors in the second title feature set to obtain several clustering sets, and then traverse the optimized feature vectors in the second candidate title feature set. If it matches the central optimized feature vector in a certain clustering set, then store it in that clustering set. In this way, several clustering sets are updated, and then the updated several clustering sets are merged to obtain the final second target title feature set, and determine the title level of the text line corresponding to the optimized feature vector in the second target title feature set.
[0307] Based on this, for the above step S210, the method shown in Figure 7 can be referred to. As shown in Figure 7 is a schematic flowchart of the implementation process of a method for determining the title level of a document title without serial number characters provided by the embodiment of the present application. This method is applied to an electronic device and specifically may include the following steps:
[0308] S701. Cluster the optimized feature vectors in the second title feature set to obtain multiple second clustering sets.
[0309] In an embodiment of the present application, for the obtained second title feature set vecNoSerialHeadings, the optimized feature vectors in the second title feature set vecNoSerialHeadings can be clustered to obtain multiple second clustering sets. Among them, the clustering algorithm used can be the kmeans algorithm, and the embodiments of the present application do not limit this.
[0310] Among them, for the optimized feature vectors in the second title feature set vecNoSerialHeadings, which include bold feature, font size feature, and document boundary feature, the optimized feature vectors in the second title feature set vecNoSerialHeadings can be clustered according to the bold feature, font size feature, and document boundary feature to obtain multiple second clustering sets.
[0311] S702. Traverse the optimized feature vectors in the second candidate title feature set, and determine the second similarity between the traversed optimized feature vector and the central optimized feature vector in any second clustering set.
[0312] In an embodiment of the present application, for the second candidate title feature set vecNoSerialDistrustHeadings, traverse the optimized feature vectors in the second candidate title feature set vecNoSerialDistrustHeadings, and for the traversed optimized feature vector, determine the second similarity between it and the central optimized feature vector in any second clustering set.
[0313] It should be noted that in the process of calculating the second similarity between the optimized feature vector and the central optimized feature vector, the differences between the same type of font size feature, document boundary feature, and bold feature can be calculated and then weighted and summed, which means that the optimized feature vectors with similar font size features, similar document boundary features, and the same font boldness as those included in the central optimized feature vector should be selected as the second candidate optimized feature vectors.
[0314] S703. Search for the largest second target similarity from the second similarities, and determine whether the largest second target similarity is greater than the preset second similarity threshold.
[0315] S704. In the case where the largest second target similarity is greater than the preset second similarity threshold, determine the traversed optimized feature vector as the second candidate optimized feature vector.
[0316] In the embodiment of the present application, through the above steps, multiple second similarity degrees can be obtained. The maximum second similarity degree can be searched from these second similarity degrees, and it is determined whether the maximum second target similarity degree is greater than a preset second similarity threshold. When the maximum second target similarity degree is greater than the preset second similarity threshold, the traversed optimized feature vector is determined as the second candidate optimized feature vector, indicating that the document title corresponding to the text behavior of the traversed optimized feature vector is missed detected, and the difference is that the text line does not carry serial number characters.
[0317] Among them, for the second candidate optimized feature vector, the second candidate optimized feature vector can be stored in the second title feature set to obtain the second target title feature set, and the title level of the text line corresponding to the optimized feature vector in the second target title feature set is determined. Specifically, after the traversed optimized feature vector is determined as the second candidate optimized feature vector, the second clustering set corresponding to the maximum second target similarity degree is searched, the second candidate optimized feature vector is stored in the searched second clustering set, and the updated second clustering sets are merged to obtain the second target title feature set. To determine the title level of the text line corresponding to the optimized feature vector in the second target title feature set, reference can be made to the subsequent steps S705 - S706.
[0318] S705, Search for the second clustering set corresponding to the maximum second target similarity degree, and store the second candidate optimized feature vector in the searched second clustering set.
[0319] In the embodiment of the present application, for the maximum second target similarity degree obtained in the above steps, the second clustering set corresponding to the maximum second target similarity degree can be searched, and the second candidate optimized feature vector is stored in the searched second clustering set, thus completing the supplement of the missed detected document title.
[0320] S706, Merge the updated second clustering sets to obtain the second target title feature set, and determine the title level of the text line corresponding to the optimized feature vector in the second target title feature set.
[0321] In the embodiment of the present application, for each second clustering set, after the above missed detected document title supplement step, an update is obtained. Thus, the updated second clustering sets are merged to obtain the second target title feature set, and the title level of the text line corresponding to the optimized feature vector in the second target title feature set is determined.
[0322] Among them, for the second target heading feature set vecAllNoSerialHeadings, it can be merged with the first target heading feature set vecAllSerialHeadings to obtain the final third target heading feature set vecAllHeadings. Sort the optimized feature vectors in the third target heading feature set vecAllHeadings to obtain the third sorting result. Traverse the optimized feature vectors in the second target heading feature set vecAllNoSerialHeadings, and find the position index corresponding to the traversed optimized feature vector in the third sorting result. According to the position index, determine the heading level of the text line corresponding to the traversed optimized feature vector.
[0323] It should be noted that for the optimized feature vectors in the third target heading feature set vecAllHeadings, they can be sorted according to the page index feature and line index feature included in the optimized feature vectors. Among them, the sorting is preferably in ascending order according to the page index feature. If the page index features are the same during this period, then the sorting is in ascending order according to the line index feature.
[0324] Specifically, with the position index as the center, forward search in the third sorting result for the nearest neighbor forward optimized feature vector vecAllHeadings[m] with the serial number feature being the serial number value, and backward search for the nearest neighbor backward optimized feature vector vecAllHeadings[n] with the serial number feature being the serial number value; according to the nearest neighbor forward optimized feature vector and the nearest neighbor backward optimized feature vector, determine the heading level of the text line corresponding to the traversed optimized feature vector.
[0325] Among them, for the nearest neighbor forward optimized feature vector and the nearest neighbor backward optimized feature vector, it is possible to find them, or it is also possible not to find them. Therefore, if the nearest neighbor forward optimized feature vector vecAllHeadings[m] is not found, then m can be set to -1. If the nearest neighbor backward optimized feature vector vecAllHeadings[n] is not found, then n can be set to vecAllHeadings.size (that is, the number of optimized feature vectors in the third target heading feature set vecAllHeadings).
[0326] Therefore, in the case where m = -1 and the nearest neighbor backward optimized feature vector vecAllHeadings[n] is found, that is, in the case where there is no nearest neighbor forward optimized feature vector but there is a nearest neighbor backward optimized feature vector, it is judged whether the traversed optimized feature vector is the first optimized feature vector in the third sorting result, that is, it is judged whether the above position index is the first position (for example, whether the position index i = 0).
[0327] When the traversed optimized feature vector is the optimized feature vector ranked first in the third sorting result, determine whether the first ratio (e.g., vecAllHeadings[i].fontsize / vecAllHeadings[n].fontsize) between the font size feature in the traversed optimized feature vector and the font size feature in the nearest neighbor backward optimized feature vector is greater than the preset second font size threshold (e.g., the preset second font size threshold is 1.2); when the first ratio is greater than the preset second font size threshold, determine that the title level of the text line corresponding to the traversed optimized feature vector vecAllHeadings[i] is the first-level title, and update the title level of the text line corresponding to the optimized feature vector with the serial number feature being the serial number value (e.g., for the title level of the text line corresponding to the optimized feature vector with the serial number feature being the serial number value, +1 on the original basis); when the first ratio is not greater than the preset second font size threshold, determine the title level of the text line corresponding to the traversed optimized feature vector as the title level corresponding to the nearest neighbor backward optimized feature vector, which means setting the title level of the text line corresponding to the traversed optimized feature vector to be the same as the title level corresponding to the nearest neighbor backward optimized feature vector.
[0328] When the traversed optimized feature vector is not the optimized feature vector ranked first in the third sorting result (e.g., the position index i is not 0), find the previous optimized feature vector of the traversed optimized feature vector in the third sorting result; determine whether the second ratio (e.g., vecAllHeadings[i - 1].fontsize / vecAllHeadings[i].fontsize) between the font size feature in the previous optimized feature vector and the font size feature in the traversed optimized feature vector is greater than the preset third font size threshold (e.g., the preset third font size threshold is 1.2); when the second ratio is greater than the preset third font size threshold, determine the title level of the text line corresponding to the traversed optimized feature vector as the sum of the title level corresponding to the previous optimized feature vector and the preset level threshold (e.g., the preset level threshold is 1); when the second ratio is not greater than the preset third font size threshold, determine the title level of the text line corresponding to the traversed optimized feature vector as the title level corresponding to the previous optimized feature vector.
[0329] In addition, if there is a nearest neighbor forward optimized feature vector vecAllHeadings[m], but n = vecAllHeadings.size (i.e., the number of optimized feature vectors in the third target heading feature set vecAllHeadings), that is, in the case where there is a nearest neighbor forward optimized feature vector but no nearest neighbor backward optimized feature vector, in the third sorting result, determine whether the nearest neighbor forward optimized feature vector is the previous optimized feature vector of the position index (for example, determine whether the position index i - 1 is the same as m); in the case where the nearest neighbor forward optimized feature vector is the previous optimized feature vector of the position index, determine whether the third ratio (for example, vecAllHeadings[i].fontsize / vecAllHeadings[m].fontsize) between the font size feature in the traversed optimized feature vector and the font size feature in the nearest neighbor forward optimized feature vector is greater than the preset fourth font size threshold (for example, the preset fourth font size threshold is 1.5); in the case where the third ratio is greater than the preset fourth font size threshold, determine whether the text line corresponding to the traversed optimized feature vector is centered according to the document boundary feature in the traversed optimized feature vector; in the case where the text line corresponding to the traversed optimized feature vector is centered, determine that the title level of the text line corresponding to the traversed optimized feature vector is a first-level title, and update the title level of the text line corresponding to the optimized feature vector with the serial number feature as the serial number value.
[0330] Alternatively, in the case where the nearest neighbor forward optimized feature vector is the previous optimized feature vector of the position index, determine whether the fourth ratio (for example, vecAllHeadings[i].fontsize / vecAllHeadings[m].fontsize) between the font size feature in the traversed optimized feature vector and the font size feature in the nearest neighbor forward optimized feature vector is less than the preset fifth font size threshold (for example, the preset fifth font size threshold is 0.9); in the case where the fourth ratio is less than the preset fifth font size threshold, determine the title level of the text line corresponding to the traversed optimized feature vector as the sum of the title level corresponding to the nearest neighbor forward optimized feature vector and the preset level threshold (for example, the preset level threshold is 1); in the case where the fourth ratio between the font size feature in the traversed optimized feature vector and the font size feature in the nearest neighbor forward optimized feature vector is not less than the preset fifth font size threshold, determine the title level of the text line corresponding to the traversed optimized feature vector as the title level corresponding to the nearest neighbor forward optimized feature vector.
[0331] In the case where the nearest neighbor forward optimized feature vector is not the previous optimized feature vector of the position index, determine whether the fifth ratio (e.g., vecAllHeadings[i].fontsize / vecAllHeadings[i - 1].fontsize) between the font size feature in the traversed optimized feature vector and the font size feature in the previous optimized feature vector is less than the preset sixth font size threshold (e.g., the preset sixth font size threshold is 0.9); in the case where the fifth ratio is less than the preset sixth font size threshold, determine the heading level of the text line corresponding to the traversed optimized feature vector as the sum of the heading level corresponding to the previous optimized feature vector and the preset level threshold; in the case where the fifth ratio is not less than the preset sixth font size threshold, determine the heading level of the text line corresponding to the traversed optimized feature vector as the heading level corresponding to the previous optimized feature vector.
[0332] It should be noted that for the first target heading feature set, the serial number feature included in the optimized feature vector therein is a serial number value (i.e., a character with a serial number), while for the second target heading feature set, the serial number feature included in the optimized feature vector therein is a second value (i.e., a character without a serial number). Merging and sorting the first target heading feature set and the second target heading feature set to obtain a third sorting result means that the third sorting result mixes optimized feature vectors with characters having serial numbers and optimized feature vectors without characters having serial numbers. Traverse the optimized feature vectors (i.e., optimized feature vectors without characters having serial numbers) in the second target heading feature set, find the corresponding position index in the third sorting result, and forward search for the nearest neighbor forward optimized feature vector with a serial number feature being a serial number value in the third sorting result centered on the position index. Since the third sorting result mixes optimized feature vectors with characters having serial numbers and optimized feature vectors without characters having serial numbers, the nearest neighbor forward optimized feature vector with a serial number feature being a serial number value may be the previous optimized feature vector of the position index (the position index corresponds to the optimized feature vector without characters having serial numbers), or may not be the previous optimized feature vector of the position index.
[0333] In addition, when both the above-mentioned m and n are legal, that is, when there is a nearest neighbor forward optimization feature vector and there is a nearest neighbor backward optimization feature vector, determine whether the reference feature vector corresponding to the nearest neighbor forward optimization feature vector is the same as the reference feature vector corresponding to the nearest neighbor backward optimization feature vector; in the case of being the same, determine the sixth ratio (for example, vecAllHeadings[i].fontsize / vecAllHeadings[m].fontsize) between the font size feature in the traversed optimization feature vector and the font size feature in the nearest neighbor forward optimization feature vector; determine the magnitude between the serial number feature in the nearest neighbor forward optimization feature vector and the serial number feature in the nearest neighbor backward optimization feature vector; in the case where the sixth ratio is greater than the preset seventh font size threshold (for example, the preset seventh font size threshold is 1.5) and the serial number feature in the nearest neighbor forward optimization feature vector is less than the serial number feature in the nearest neighbor backward optimization feature vector (indicating that the serial number characters are in descending order, for example, between (1) and (3) is in descending order), determine the heading level of the text line corresponding to the traversed optimization feature vector as a first-level heading, and update the heading level of the text line corresponding to the optimization feature vector with the serial number feature being the serial number value; in the case where the sixth ratio is not greater than the preset seventh font size threshold and / or the serial number feature in the nearest neighbor forward optimization feature vector is not less than the serial number feature in the nearest neighbor backward optimization feature vector, determine the heading level of the text line corresponding to the traversed optimization feature vector as the sum of the heading level corresponding to the nearest neighbor forward optimization feature vector and the preset level threshold.
[0334] In the case where the reference feature vector corresponding to the nearest neighbor forward optimization feature vector is different from the reference feature vector corresponding to the nearest neighbor backward optimization feature vector, determine whether the heading level corresponding to the nearest neighbor forward optimization feature vector is greater than the heading level corresponding to the nearest neighbor backward optimization feature vector; in the case where the heading level corresponding to the nearest neighbor forward optimization feature vector is greater than the heading level corresponding to the nearest neighbor backward optimization feature vector, in the third sorting result, starting from the nearest neighbor backward optimization feature vector, search backward for a new nearest neighbor backward optimization feature vector, and jump to the processing step of the case where the reference feature vector corresponding to the nearest neighbor forward optimization feature vector is the same as the reference feature vector corresponding to the nearest neighbor backward optimization feature vector, where the reference feature vector of the new nearest neighbor backward optimization feature vector is the same as the reference feature vector corresponding to the nearest neighbor forward optimization feature vector;
[0335] In the case where the title level corresponding to the nearest neighbor forward optimization feature vector is less than the title level corresponding to the nearest neighbor backward optimization feature vector, in the third sorting result, starting from the nearest neighbor forward optimization feature vector, search forward for a new nearest neighbor forward optimization feature vector, and jump to the processing step in the case where the reference feature vector corresponding to the nearest neighbor forward optimization feature vector is the same as the reference feature vector corresponding to the nearest neighbor backward optimization feature vector, where the reference feature vector of the new nearest neighbor forward optimization feature vector is the same as the reference feature vector corresponding to the nearest neighbor backward optimization feature vector; in the case where the title level corresponding to the nearest neighbor forward optimization feature vector is equal to the title level corresponding to the nearest neighbor backward optimization feature vector, determine the title level of the text line corresponding to the traversed optimization feature vector as the sum of the title level corresponding to the nearest neighbor forward optimization feature vector and the preset level threshold.
[0336] After going through the above steps, for each optimization feature vector in the third target title feature set, there is a corresponding title level for its corresponding text line (i.e., the document title), and a reasonably structured title level can be obtained.
[0337] Corresponding to the above method embodiment, an embodiment of the present application also provides a device for determining the document title level, which may include: a text line acquisition module, a vector extraction module, a title determination module, and a title level determination module.
[0338] The text line acquisition module is used to acquire a target document and acquire the text lines in the target document page of the target document, and store them in the title candidate line set;
[0339] The vector extraction module is used to extract the basic feature vector and the optimization feature vector of each text line in the title candidate line set;
[0340] The title determination module is used to determine the document title in the title candidate line set according to the basic feature vector and the optimization feature vector;
[0341] The title level determination module is used to determine the title level of the document title.
[0342] In an optional embodiment, the vector extraction module specifically includes:
[0343] The feature extraction sub-module is used to extract the optimization features of each text line in the title candidate line set;
[0344] The feature combination sub-module is used to form an optimization feature vector from the optimization features; where the optimization features include at least one of the following: font application feature, bold feature, serial number feature, font size feature, document boundary feature, page index feature, line index feature, text feature, color feature.
[0345] In an alternative embodiment, when the optimization feature is a font application feature, the feature extraction sub-module is specifically configured to: for any text line in the set of title candidate lines, obtain the fonts corresponding to each character in the text line; when all the fonts are the same, set the font application feature to a first value; when all or some of the fonts are different, set the font application feature to a second value;
[0346] Alternatively, when the optimization feature is a bold feature, the feature extraction sub-module is specifically configured to: for any text line in the set of title candidate lines, obtain the fonts corresponding to each character in the text line; when all the fonts are the same and are bold fonts, set the bold feature to a first value;
[0347] Or,
[0348] for any text line in the set of title candidate lines, obtain the fonts corresponding to each character in the text line and the font weights; when all the fonts are the same and all the font weights are greater than a preset weight threshold, set the bold feature to a first value;
[0349] Or,
[0350] for any text line in the set of title candidate lines, obtain the fonts corresponding to each character in the text line and the font weights; when all or some of the fonts are different, all the font weights are not greater than the preset weight threshold, and all the fonts are not bold fonts, perform the following processing; classify the characters included in the text line according to the font types included in the text line, and set the bold feature for the characters included in the text line for different font types; wherein, classifying the characters included in the text line according to the font types included in the text line and setting the bold feature for the characters included in the text line for different font types includes: rendering the text line by font to obtain at least two third text line images, and searching for the next text line of the text line from the set of title candidate lines; rendering the next text line to obtain a second text line image, and setting the bold feature according to at least two of the third text line images and the second text line image.
[0351] In an alternative embodiment, the feature extraction sub-module is further configured to:
[0352] when all the fonts are the same, all the font weights are not greater than the preset weight threshold, and all the fonts are not bold fonts, perform the following processing;
[0353] Render the text line to obtain a first text line image, and find the next text line of the text line from the set of title candidate lines;
[0354] Render the next text line to obtain a second text line image, and set the bold feature according to the first text line image and the second text line image.
[0355] In an alternative embodiment, the feature extraction sub-module is further configured to:
[0356] Input the first text line image and the second text line image into a pre-trained bold classification model to obtain a first output result;
[0357] In the case where the first output result is a preset result, set the bold feature to a first value;
[0358] Alternatively, in the case where the first output result is not a preset result, set the bold feature to a second value.
[0359] In an alternative embodiment, the feature extraction sub-module is further configured to:
[0360] For any one of the third text line images, input the third text line image and the second text line image into a pre-trained bold classification model to obtain a second output result;
[0361] In the case where all the second output results are preset results, set the bold feature to a first value;
[0362] Alternatively, in the case where all or some of the second output results are not preset results, set the bold feature to a second value.
[0363] In an alternative embodiment, the apparatus further includes: a bold classification model training module, configured to obtain a character image pair and a sample label corresponding to the character image pair, input the character image pair into the bold classification model to obtain a prediction result; determine a loss value between the prediction result and the sample label, and train the bold classification model according to the loss value; stop training when the loss value converges to obtain a pre-trained bold classification model.
[0364] In an alternative embodiment, when the optimization feature is a serial number feature, the feature extraction sub-module is specifically configured to: for any text line in the set of title candidate lines, detect whether there is a serial number character at the beginning of the text line; in the case where there is a serial number character, find the serial number value corresponding to the serial number character, and set the serial number feature to the serial number value; or, in the case where there is no serial number character, set the serial number feature to a second value;
[0365] Alternatively, when the optimization feature is the font size feature, the feature extraction sub-module is specifically configured to: for any text line in the set of candidate title lines, obtain the font sizes corresponding to each character in the text line; determine the average font size of the font sizes, and set the font size feature to the average font size;
[0366] Alternatively, when the optimization feature is the document boundary feature, the feature extraction sub-module is specifically configured to: for any text line in the set of candidate title lines, determine the distance between the text line and the document boundary, and set the document boundary feature to the distance;
[0367] Alternatively, when the optimization features are the page index feature and the line index feature, the feature extraction sub-module is specifically configured to: for any text line in the set of candidate title lines, determine the page index and the line index of the text line; set the page index feature to the page index, and set the line index feature to the line index;
[0368] Alternatively, when the optimization feature is the text feature, the feature extraction sub-module is specifically configured to: for any text line in the set of candidate title lines, obtain the characters included in the text line, and set the text feature to the characters included in the text line;
[0369] Alternatively, when the optimization feature is the color feature, the feature extraction sub-module is specifically configured to: for any text line in the set of candidate title lines, determine the character area in the text line; determine the average pixel value of the character area, and set the color feature to the average pixel value.
[0370] In an alternative embodiment, the title determination module is specifically configured to:
[0371] For any text line in the set of candidate title lines, input the basic feature vector and the optimization feature vector of the text line into a pre-trained title classifier to obtain a classification probability value;
[0372] When the classification probability value is greater than a preset classification threshold, determine the text line as the document title, and store the optimization feature vector of the text line in the title feature set;
[0373] When the classification probability value is not greater than the preset classification threshold, store the optimization feature vectors of the text lines that meet the first preset condition in the candidate title feature set; determine the title level of the document title according to the title feature set and the candidate title feature set.
[0374] In an alternative embodiment, the device further includes: a title classifier training module, configured to obtain a sample title corpus, and extract a sample basic feature vector and a sample optimized feature vector of the sample title corpus; input the sample basic feature vector and the sample optimized feature vector into a title classifier to obtain a predicted classification probability value; determine a probability loss between the predicted classification probability value and a sample classification probability value corresponding to the sample title corpus; train the title classifier according to the probability loss, and stop training when the probability loss converges to obtain a pre-trained title classifier.
[0375] In an alternative embodiment, the title level determination module specifically includes:
[0376] A title feature set division sub-module, configured to divide the title feature set into a first title feature set and a second title feature set;
[0377] A candidate title feature set division sub-module, configured to divide the candidate title feature set into a first candidate title feature set and a second candidate title feature set;
[0378] A first title level determination sub-module, configured to determine a first candidate optimized feature vector in the first candidate title feature set, store it in the first title feature set to obtain a first target title feature set, and determine the title level of the text line corresponding to the optimized feature vector in the first target title feature set;
[0379] A second title level determination sub-module, configured to determine a second candidate optimized feature vector in the second candidate title feature set, store it in the second title feature set to obtain a second target title feature set, and determine the title level of the text line corresponding to the optimized feature vector in the second target title feature set.
[0380] In an alternative embodiment, the title feature set division sub-module is specifically configured to:
[0381] Traverse the optimized feature vectors of the title feature set to obtain the serial number features in the traversed optimized feature vectors;
[0382] If the serial number feature is not a second value, store the optimized feature vector in the first title feature set;
[0383] If the serial number feature is the second value, store the optimized feature vector in the second title feature set;
[0384] The candidate title feature set division sub-module is specifically configured to:
[0385] Traverse the optimized feature vectors in the candidate title feature set, and obtain the serial number features in the traversed optimized feature vectors;
[0386] When the serial number feature is not the second value, store the optimized feature vector into the first candidate title feature set;
[0387] When the serial number feature is the second value, store the optimized feature vector into the second candidate title feature set.
[0388] In an optional implementation manner, the first title level determination sub-module specifically includes:
[0389] The first clustering unit is used to cluster the optimized feature vectors in the first title feature set to obtain a plurality of first clustering sets;
[0390] The benchmark feature vector determination unit is used to determine the benchmark feature vector corresponding to each first clustering set;
[0391] The first similarity determination unit is used to traverse the optimized feature vectors in the first candidate title feature set, and determine the first similarity between the traversed optimized feature vector and any one of the benchmark feature vectors;
[0392] The first target similarity judgment unit is used to find the largest first target similarity from the first similarities, and judge whether the first target similarity is greater than a preset first similarity threshold;
[0393] The first candidate optimized feature vector determination unit is used to determine the traversed optimized feature vector as the first candidate optimized feature vector when the first target similarity is greater than the preset first similarity threshold;
[0394] The first title level determination unit is used to store the first candidate optimized feature vector into the first title feature set to obtain a first target title feature set, and determine the title level of the text line corresponding to the optimized feature vector in the first target title feature set;
[0395] Among them, the first title level determination unit specifically includes:
[0396] The first clustering set update sub-unit is used to determine the benchmark feature vector corresponding to the first target similarity, find the first clustering set corresponding to the determined benchmark feature vector, and store the first candidate optimized feature vector into the found first clustering set;
[0397] The first title level determination subunit is used to determine the title level of the text lines corresponding to the optimized feature vectors in each updated first clustering set; wherein, each updated first clustering set is merged to obtain a first target title feature set.
[0398] In an optional implementation manner, the first clustering unit is specifically configured to:
[0399] Obtain the serial number features included in the optimized feature vectors in the first title feature set, and determine whether the serial number features are the same;
[0400] In the case where the serial number features are different, cluster the optimized feature vectors in the first title feature set according to the serial number features to obtain multiple first clustering sets;
[0401] Alternatively, in the case where the serial number features are the same, cluster the optimized feature vectors in the first title feature set according to the font size features to obtain multiple first clustering sets.
[0402] In an optional implementation manner, the reference feature vector determination unit specifically includes:
[0403] A reference feature extraction subunit, configured to extract the reference features corresponding to any one of the first clustering sets;
[0404] A reference feature combination subunit, configured to form a reference feature vector from the reference features; wherein, the reference features include at least one of the following: reference font application feature, reference bold feature, reference serial number feature, reference font size feature, reference document boundary feature, reference page index feature, reference line index feature, reference text feature, reference color feature.
[0405] In an optional implementation manner, when the reference feature is a reference font application feature, the reference feature extraction subunit is specifically configured to: count the first font application quantity corresponding to the optimized feature vectors with the font application feature being the first value in the first clustering set; count the second font application quantity corresponding to the optimized feature vectors with the font application feature being the second value in the first clustering set; in the case where the first font application quantity is greater than the second font application quantity, set the reference font application feature corresponding to the first clustering set to the first value; or, in the case where the first font application quantity is not greater than the second font application quantity, set the reference font application feature corresponding to the first clustering set to the second value;
[0406] Alternatively, when the reference feature is a reference bold feature, the reference feature extraction subunit is specifically configured to: count the first bold quantity corresponding to the optimized feature vectors with the bold feature being the first value in the first clustering set; count the second non-bold quantity corresponding to the optimized feature vectors with the bold feature being the second value in the first clustering set; when the first bold quantity is greater than the second non-bold quantity, set the reference bold feature corresponding to the first clustering set to the first value; or, when the first bold quantity is not greater than the second non-bold quantity, set the reference bold feature corresponding to the first clustering set to the second value;
[0407] Alternatively, when the reference feature is a reference serial number feature, the reference feature extraction subunit is specifically configured to: obtain the serial number feature included in the optimized feature vectors in the first clustering set, and set the reference serial number feature corresponding to the first clustering set to the serial number feature;
[0408] Alternatively, when the reference feature is a reference font size feature, the reference feature extraction subunit is specifically configured to: obtain the font size feature included in the optimized feature vectors in the first clustering set, and determine the average font size feature of the font size feature; set the reference font size feature corresponding to the first clustering set to the average font size feature;
[0409] Alternatively, when the reference feature is a reference document boundary feature, the reference feature extraction subunit is specifically configured to: obtain the document boundary feature included in the optimized feature vectors in the first clustering set, and determine the average document boundary feature of the document boundary feature;
[0410] Set the reference document boundary feature corresponding to the first clustering set to the average document boundary feature;
[0411] Alternatively, when the reference feature is a reference page index feature or a reference line index feature, the reference feature extraction subunit is specifically configured to: obtain the page index feature included in the optimized feature vectors in the first clustering set, and select the smallest page index feature from the page index features; set the reference page index feature corresponding to the first clustering set to the smallest page index feature; obtain the line index feature included in the optimized feature vectors in the first clustering set, and select the smallest line index feature from the line index features; set the reference line index feature corresponding to the first clustering set to the smallest line index feature;
[0412] Alternatively, when the reference feature is a reference text feature, the reference feature extraction subunit is specifically configured to: set the reference text feature corresponding to the first clustering set to empty, where empty means not filling in any content;
[0413] Alternatively, when the reference feature is a reference color feature, the reference feature extraction subunit is specifically configured to: obtain the color features included in the optimized feature vectors in the first clustering set, and determine the average color feature of the color features; set the reference color feature corresponding to the first clustering set to the average color feature.
[0414] In an alternative embodiment, the first title level determination subunit is specifically configured to:
[0415] Sort all the reference feature vectors to obtain a first sorting result;
[0416] According to the first sorting result, determine the title levels of the text lines corresponding to the optimized feature vectors in each updated first clustering set;
[0417] Wherein, the first title level determination subunit is specifically configured to:
[0418] Traverse the reference feature vectors in the first sorting result in sequence:
[0419] When the traversed reference feature vector is the first reference feature vector, find the updated first clustering set corresponding to the traversed reference feature vector; determine the first-level title corresponding to the traversed reference feature vector as the title level of the text line corresponding to the optimized feature vector in the updated first clustering set corresponding to the traversed reference feature vector;
[0420] Alternatively, when the traversed reference feature vector is not the first reference feature vector, find the previous reference feature vector of the traversed reference feature vector from the first sorting result; according to the previous reference feature vector, determine the title level of the text line corresponding to the optimized feature vector in the updated first clustering set corresponding to the traversed reference feature vector.
[0421] In an alternative embodiment, the first title level determination subunit is further configured to:
[0422] When the reference page index feature included in the previous reference feature vector is the same as the reference page index feature included in the traversed reference feature vector, determine the previous title level corresponding to the previous reference feature vector;
[0423] Determine the sum of the previous title level and a preset level threshold as the title level corresponding to the traversed reference feature vector;
[0424] Determine the title level corresponding to the traversed reference feature vector as the title level of the text line corresponding to the optimized feature vector in the updated first clustering set corresponding to the traversed reference feature vector;
[0425] Add the traversed reference feature vector to the subtitle list of the previous reference feature vector;
[0426] Alternatively, when the reference page index feature included in the previous reference feature vector is different from the reference page index feature included in the traversed reference feature vector, sort the optimized feature vectors in each updated first clustering set to obtain a second sorting result;
[0427] For the updated first clustering set corresponding to the traversed reference feature vector, traverse the optimized feature vectors in the updated first clustering set, and find the previous optimized feature vector of the traversed optimized feature vector from the second sorting result;
[0428] In the case where there is no previous optimized feature vector, find the next optimized feature vector of the traversed optimized feature vector from the second sorting result; or, in the case where there is a next optimized feature vector, find the reference feature vector corresponding to the updated first clustering set where the next optimized feature vector is located;
[0429] When the traversed reference feature vector is different from the reference feature vector corresponding to the updated first clustering set where the next optimized feature vector is located, and the product of the reference font size feature included in the next optimized feature vector and a preset first multiple is less than the reference font size feature included in the traversed reference feature vector, determine the next title level of the reference feature vector corresponding to the updated first clustering set where the next optimized feature vector is located;
[0430] Determine the difference between the next title level and the preset level threshold as the title level corresponding to the traversed reference feature vector; determine the title level corresponding to the traversed reference feature vector as the title level of the text line corresponding to the optimized feature vector in the updated first clustering set corresponding to the traversed reference feature vector;
[0431] Add the reference feature vector corresponding to the updated first clustering set where the next optimized feature vector is located to the subtitle list of the traversed reference feature vector.
[0432] In an alternative embodiment, the first title level determination subunit is further configured to:
[0433] In the case where there is a previous optimized feature vector, determine the previous title level of the reference feature vector corresponding to the updated first clustering set where the previous optimized feature vector is located;
[0434] Determine the sum of the previous title level and the preset level threshold as the title level corresponding to the traversed reference feature vector;
[0435] Determine the title level corresponding to the traversed reference feature vector as the title level of the text line corresponding to the optimized feature vector in the updated first clustering set corresponding to the traversed reference feature vector;
[0436] Add the traversed reference feature vector to the sub - title list of the subsequent reference feature vector.
[0437] In an alternative embodiment, the second title level determination sub - module specifically includes:
[0438] A second clustering unit, configured to cluster the optimized feature vectors in the second title feature set to obtain multiple second clustering sets;
[0439] A second similarity determination unit, configured to traverse the optimized feature vectors in the second candidate title feature set, and determine the second similarity between the traversed optimized feature vector and the central optimized feature vector in any second clustering set;
[0440] A second target similarity judgment unit, configured to find the maximum second target similarity from the second similarities, and judge whether the maximum second target similarity is greater than a preset second similarity threshold;
[0441] A second candidate optimized feature vector determination unit, configured to, when the maximum second target similarity is greater than the preset second similarity threshold, determine the traversed optimized feature vector as the second candidate optimized feature vector;
[0442] A second title level determination unit, configured to store the second candidate optimized feature vector into the second title feature set to obtain a second target title feature set, and determine the title level of the text line corresponding to the optimized feature vector in the second target title feature set;
[0443] Among them, the second title level determination unit specifically includes:
[0444] A second clustering set update sub - unit, configured to find the second clustering set corresponding to the maximum second target similarity, and store the second candidate optimized feature vector into the found second clustering set;
[0445] A second title level determination sub - unit, configured to merge the updated second clustering sets to obtain a second target title feature set, and determine the title level of the text line corresponding to the optimized feature vector in the second target title feature set.
[0446] In an alternative embodiment, the second clustering unit is specifically configured to:
[0447] Cluster the optimized feature vectors in the second title feature set according to the bold feature, font size feature, and document boundary feature to obtain multiple second clustering sets.
[0448] In an optional embodiment, the second title level determination subunit is specifically configured to:
[0449] Merge the first target title feature set and the second target title feature set to obtain a third target title feature set;
[0450] Sort the optimized feature vectors in the third target title feature set to obtain a third sorting result;
[0451] Traverse the optimized feature vectors in the second target title feature set, and find the position index corresponding to the traversed optimized feature vector in the third sorting result;
[0452] Determine the title level of the text line corresponding to the traversed optimized feature vector according to the position index.
[0453] In an optional embodiment, the second title level determination subunit is specifically configured to:
[0454] Centering on the position index, forward search in the third sorting result for the nearest neighbor forward optimized feature vector with the serial number feature being the serial number value, and backward search for the nearest neighbor backward optimized feature vector with the serial number feature being the serial number value;
[0455] Determine the title level of the text line corresponding to the traversed optimized feature vector according to the nearest neighbor forward optimized feature vector and the nearest neighbor backward optimized feature vector.
[0456] In an optional embodiment, the second title level determination subunit is specifically configured to:
[0457] In the case where there is no nearest neighbor forward optimized feature vector but there is a nearest neighbor backward optimized feature vector, determine whether the traversed optimized feature vector is the first optimized feature vector in the third sorting result;
[0458] In the case where the traversed optimized feature vector is the first optimized feature vector in the third sorting result, determine whether the first ratio between the font size feature in the traversed optimized feature vector and the font size feature in the nearest neighbor backward optimized feature vector is greater than a preset second font size threshold; in the case where the first ratio is greater than the preset second font size threshold, determine that the title level of the text line corresponding to the traversed optimized feature vector is the first-level title, and update the title level of the text line corresponding to the optimized feature vector with the serial number feature being the serial number value;
[0459] Alternatively, when the first ratio is not greater than a preset second font size threshold, determine the title level of the text line corresponding to the traversed optimized feature vector as the title level corresponding to the nearest neighbor backward optimized feature vector;
[0460] Or,
[0461] When the traversed optimized feature vector is not the optimized feature vector ranked first in the third sorting result, find the previous optimized feature vector of the traversed optimized feature vector in the third sorting result; determine whether the second ratio between the font size feature in the previous optimized feature vector and the font size feature in the traversed optimized feature vector is greater than a preset third font size threshold; when the second ratio is greater than the preset third font size threshold, determine the title level of the text line corresponding to the traversed optimized feature vector as the sum of the title level corresponding to the previous optimized feature vector and a preset level threshold; or, when the second ratio is not greater than the preset third font size threshold, determine the title level of the text line corresponding to the traversed optimized feature vector as the title level corresponding to the previous optimized feature vector.
[0462] In an alternative embodiment, the second title level determination subunit is further configured to:
[0463] When there is a nearest neighbor forward optimized feature vector but no nearest neighbor backward optimized feature vector, in the third sorting result, determine whether the nearest neighbor forward optimized feature vector is the previous optimized feature vector of the position index;
[0464] When the nearest neighbor forward optimized feature vector is the previous optimized feature vector of the position index, determine whether the third ratio between the font size feature in the traversed optimized feature vector and the font size feature in the nearest neighbor forward optimized feature vector is greater than a preset fourth font size threshold; when the third ratio is greater than the preset fourth font size threshold, according to the document boundary feature in the traversed optimized feature vector, determine whether the text line corresponding to the traversed optimized feature vector is centered; when the text line corresponding to the traversed optimized feature vector is centered, determine the title level of the text line corresponding to the traversed optimized feature vector as the first-level title, and update the title level of the text line corresponding to the optimized feature vector with the serial number feature being the serial number value;
[0465] Or,
[0466] When the nearest neighbor forward optimized feature vector is the previous optimized feature vector of the position index, determine whether the fourth ratio between the font size feature in the traversed optimized feature vector and the font size feature in the nearest neighbor forward optimized feature vector is less than a preset fifth font size threshold; when the fourth ratio is less than the preset fifth font size threshold, determine the title level of the text line corresponding to the traversed optimized feature vector as the sum of the title level corresponding to the nearest neighbor forward optimized feature vector and the preset level threshold; or, when the fourth ratio between the font size feature in the traversed optimized feature vector and the font size feature in the nearest neighbor forward optimized feature vector is not less than the preset fifth font size threshold, determine the title level of the text line corresponding to the traversed optimized feature vector as the title level corresponding to the nearest neighbor forward optimized feature vector;
[0467] Or,
[0468] When the nearest neighbor forward optimized feature vector is not the previous optimized feature vector of the position index, determine whether the fifth ratio between the font size feature in the traversed optimized feature vector and the font size feature in the previous optimized feature vector is less than a preset sixth font size threshold; when the fifth ratio is less than the preset sixth font size threshold, determine the title level of the text line corresponding to the traversed optimized feature vector as the sum of the title level corresponding to the forward optimized feature vector and the preset level threshold; or, when the fifth ratio is not less than the preset sixth font size threshold, determine the title level of the text line corresponding to the traversed optimized feature vector as the title level corresponding to the forward optimized feature vector;
[0469] Or,
[0470] In the case where there is a nearest neighbor forward optimization feature vector and a nearest neighbor backward optimization feature vector, determine whether the reference feature vector corresponding to the nearest neighbor forward optimization feature vector is the same as the reference feature vector corresponding to the nearest neighbor backward optimization feature vector; in the case of being the same, determine the sixth ratio between the font size feature in the traversed optimization feature vector and the font size feature in the nearest neighbor forward optimization feature vector; determine the magnitude between the serial number feature in the nearest neighbor forward optimization feature vector and the serial number feature in the nearest neighbor backward optimization feature vector; in the case where the sixth ratio is greater than the preset seventh font size threshold and the serial number feature in the nearest neighbor forward optimization feature vector is less than the serial number feature in the nearest neighbor backward optimization feature vector, determine the title level of the text line corresponding to the traversed optimization feature vector as the first-level title, and update the title level of the text line corresponding to the optimization feature vector with the serial number feature being the serial number value; or, in the case where the sixth ratio is not greater than the preset seventh font size threshold and / or the serial number feature in the nearest neighbor forward optimization feature vector is not less than the serial number feature in the nearest neighbor backward optimization feature vector, determine the title level of the text line corresponding to the traversed optimization feature vector as the sum of the title level corresponding to the nearest neighbor forward optimization feature vector and the preset level threshold.
[0471] In an alternative embodiment, the second title level determination subunit is further configured to:
[0472] In different cases, determine whether the title level corresponding to the nearest neighbor forward optimization feature vector is greater than the title level corresponding to the nearest neighbor backward optimization feature vector; in the case where the title level corresponding to the nearest neighbor forward optimization feature vector is greater than the title level corresponding to the nearest neighbor backward optimization feature vector, in the third sorting result, starting from the nearest neighbor backward optimization feature vector, search backward for a new nearest neighbor backward optimization feature vector, and jump to the processing steps of the same case, where the reference feature vector of the new nearest neighbor backward optimization feature vector is the same as the reference feature vector corresponding to the nearest neighbor forward optimization feature vector;
[0473] Or, in the case where the title level corresponding to the nearest neighbor forward optimization feature vector is less than the title level corresponding to the nearest neighbor backward optimization feature vector, in the third sorting result, starting from the nearest neighbor forward optimization feature vector, search forward for a new nearest neighbor forward optimization feature vector, and jump to the processing steps of the same case, where the reference feature vector of the new nearest neighbor forward optimization feature vector is the same as the reference feature vector corresponding to the nearest neighbor backward optimization feature vector;
[0474] Alternatively, when the title level corresponding to the nearest neighbor forward-optimized feature vector is equal to the title level corresponding to the nearest neighbor backward-optimized feature vector, the title level of the text line corresponding to the traversed optimized feature vector is determined as the sum of the title level corresponding to the nearest neighbor forward-optimized feature vector and the preset level threshold.
[0475] An embodiment of the present application further provides an electronic device. As shown in the figure, it includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus.
[0476] The memory is used to store a computer program.
[0477] When the processor is used to execute the program stored on the memory, the following steps are implemented:
[0478] Obtain a target document, and obtain the text lines in the target document page of the target document, and store them in a title candidate line set; extract the basic feature vectors and optimized feature vectors of each text line in the title candidate line set; determine the document title in the title candidate line set according to the basic feature vectors and the optimized feature vectors, and determine the title level of the document title.
[0479] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is used in the figure to represent it, but it does not mean that there is only one bus or one type of bus.
[0480] The communication interface is used for communication between the above electronic device and other devices.
[0481] The memory may include a Random Access Memory (RAM), or may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0482] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU for short), a Network Processor (NP for short), etc.; it may also be a Digital Signal Processor (DSP for short), an Application Specific Integrated Circuit (ASIC for short), a Field-Programmable Gate Array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0483] In another embodiment provided by the present application, a storage medium is further provided. Instructions are stored in the storage medium. When it runs on a computer, the computer is enabled to execute the method for determining the document title hierarchy described in any one of the above embodiments.
[0484] In another embodiment provided by the present application, a computer program product containing instructions is further provided. When it runs on a computer, the computer is enabled to execute the method for determining the document title hierarchy described in any one of the above embodiments.
[0485] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions may be stored in a storage medium, or transmitted from one storage medium to another storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The storage medium may be any available medium that can be accessed by a computer, or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a Solid State Disk (SSD)).
[0486] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
[0487] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and reference can be made to the relevant part of the method embodiment for the related content.
[0488] The above description is only a preferred embodiment of the present application and is not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application are all included in the protection scope of the present application.
Claims
1. A method for determining a document title level, characterized in that: The method comprises: Obtain a target document, and obtain text lines in a target document page of the target document, and store them in a title candidate line set; Extracting a basic feature vector and an optimized feature vector of each text line in the title candidate line set; According to the basic feature vector and the optimized feature vector, the document title in the title candidate row set is determined, and the title level of the document title is determined.
2. The method according to claim 1, characterized in that Extracting an optimized feature vector of each text line in the title candidate line set includes: Extracting optimized features of each of the text lines in the title candidate line set, and forming an optimized feature vector from the optimized features; The optimization feature includes at least one of the following: font application feature, bold feature, serial number feature, font size feature, document boundary feature, page index feature, line index feature, text feature, and color feature.
3. The method according to claim 2, characterized in that In the case where the optimization feature is a font application feature, extracting the optimization feature of each text line in the title candidate line set includes: For any of the text lines in the title candidate line set, obtaining the font corresponding to each character in the text line; When all the fonts are the same, setting the font application characteristic to a first value; When all or part of the fonts are different, the font application feature is set to a second value.
4. The method according to claim 2, characterized in that: In the case where the optimization feature is a bold feature, extracting the optimization feature of each text line in the title candidate line set includes: For any of the text lines in the title candidate line set, obtain the font corresponding to each character in the text line; if all the fonts are the same and bold, set the bold feature to a first value; or, For any of the text lines in the title candidate line set, obtain the font and font thickness corresponding to each character in the text line; if all the fonts are the same and all the font thicknesses are greater than a preset thickness threshold, set the bold feature to a first value; or, For any of the text lines in the title candidate line set, obtain the font and font weight corresponding to each character in the text line; when all or part of the fonts are different, the weights of all the fonts are not greater than a preset weight threshold, and all the fonts are not bold fonts, perform the following processing: classify the text contained in the text line according to the font type contained in the text line, and set bold features for the text contained in the text line according to different font types; wherein, classify the text contained in the text line according to the font type contained in the text line, and set bold features for the text contained in the text line according to different font types, including: rendering the text line by font to obtain at least two third text line images, searching for the next text line of the text line from the title candidate line set; rendering the next text line to obtain a second text line image, and setting the bold feature according to the at least two third text line images and the second text line image.
5. The method according to claim 4, characterized in that The method further comprises: In the case that all the fonts are the same, the thickness of all the fonts is not greater than a preset thickness threshold, and all the fonts are not bold fonts, the following processing is performed; Rendering the text line to obtain a first text line image, and searching for a next text line of the text line from the title candidate line set; The next text line is rendered to obtain a second text line image, and the bold feature is set according to the first text line image and the second text line image.
6. The method according to claim 5, characterized in that The step of setting the bold feature according to the first text line image and the second text line image includes: Inputting the first text line image and the second text line image into a pre-trained bold classification model to obtain a first output result; When the first output result is a preset result, setting the bold feature to a first value; Alternatively, when the first output result is not a preset result, the bold feature is set to a second value.
7. The method according to claim 4, characterized in that The step of setting the bold feature according to at least two of the third text line images and the second text line image includes: For any of the third text line images, input the third text line image and the second text line image into a pre-trained bold classification model to obtain a second output result; When all the second output results are preset results, setting the bold feature to a first value; Alternatively, when all or part of the second output results are not preset results, the bold feature is set to a second value.
8. The method according to claim 2, characterized in that: In the case where the optimization feature is a sequence number feature, extracting the optimization feature of each text line in the title candidate line set includes: for any text line in the title candidate line set, detecting whether there is a sequence number character at the beginning of the text line; if there is a sequence number character, finding a sequence number value corresponding to the sequence number character, and setting the sequence number feature to the sequence number value; or, if there is no sequence number character, setting the sequence number feature to a second value; Alternatively, in the case where the optimization feature is a font size feature, the extracting the optimization feature of each text line in the title candidate line set includes: for any text line in the title candidate line set, obtaining the font size corresponding to each character in the text line; determining an average font size of the font sizes, and setting the font size feature to the average font size; Alternatively, in the case where the optimization feature is a document boundary feature, extracting the optimization feature of each text line in the title candidate line set includes: for any text line in the title candidate line set, determining the distance between the text line and the document boundary, and setting the document boundary feature to the distance; Alternatively, in the case where the optimization feature is a page index feature or a row index feature, the extracting the optimization feature of each text line in the title candidate line set includes: for any text line in the title candidate line set, determining the page index and the row index of the text line; setting the page index feature to the page index, and setting the row index feature to the row index; Alternatively, in the case where the optimization feature is a text feature, extracting the optimization feature of each text line in the title candidate line set includes: for any text line in the title candidate line set, obtaining characters contained in the text line, and setting the text feature to the characters contained in the text line; Alternatively, in the case where the optimization feature is a color feature, extracting the optimization feature of each text line in the title candidate line set includes: for any text line in the title candidate line set, determining the character area in the text line; determining the pixel average value of the character area, and setting the color feature as the pixel average value.
9. The method according to claim 1, characterized in that: The step of determining the document titles in the title candidate row set according to the basic feature vector and the optimized feature vector includes: For any of the text lines in the title candidate line set, inputting the basic feature vector and the optimized feature vector of the text line into a pre-trained title classifier to obtain a classification probability value; In the case where the classification probability value is greater than a preset classification threshold, determining the text line as a document title, and storing the optimized feature vector of the text line into a title feature set; Alternatively, when the classification probability value is not greater than a preset classification threshold, the optimized feature vector of the text line that meets the first preset condition is stored in a candidate title feature set; and the title level of the document title is determined based on the title feature set and the candidate title feature set.
10. The method according to claim 9, characterized in that Determining the title level of the document title includes: Dividing the title feature set into a first title feature set and a second title feature set; Dividing the candidate title feature set into a first candidate title feature set and a second candidate title feature set; Determine a first candidate optimized feature vector in the first candidate title feature set, store it in the first title feature set to obtain a first target title feature set, and determine the title level of the text line corresponding to the optimized feature vector in the first target title feature set; Determine a second candidate optimized feature vector in the second candidate title feature set, store it in the second title feature set to obtain a second target title feature set, and determine the title level of the text line corresponding to the optimized feature vector in the second target title feature set.
11. The method according to claim 10, characterized in that: The dividing the title feature set into a first title feature set and a second title feature set includes: Traversing the optimized feature vectors of the title feature set, and obtaining the sequence number features in the traversed optimized feature vectors; When the sequence number feature is not the second value, storing the optimized feature vector into the first title feature set; When the sequence number feature is a second value, storing the optimized feature vector into a second title feature set; The step of dividing the candidate title feature set into a first candidate title feature set and a second candidate title feature set comprises: Traversing the optimized feature vectors in the candidate title feature set, and obtaining the sequence number features in the traversed optimized feature vectors; When the sequence number feature is not the second value, storing the optimized feature vector into the first candidate title feature set; When the sequence number feature is a second value, the optimized feature vector is stored in a second candidate title feature set.
12. The method according to claim 10, characterized in that The step of determining a first candidate optimized feature vector in the first candidate title feature set, storing the first title feature set in the first title feature set to obtain a first target title feature set, and determining a title level of a text line corresponding to the optimized feature vector in the first target title feature set includes: Clustering the optimized feature vectors in the first title feature set to obtain a plurality of first cluster sets, and determining a reference feature vector corresponding to each of the first cluster sets; Traversing the optimized feature vectors in the first candidate title feature set, and determining a first similarity between the traversed optimized feature vector and any of the reference feature vectors; Finding the largest first target similarity from the first similarities, and determining whether the first target similarity is greater than a preset first similarity threshold; When the first target similarity is greater than a preset first similarity threshold, determining the traversed optimized feature vector as a first candidate optimized feature vector; Storing the first candidate optimized feature vector in the first title feature set to obtain a first target title feature set, and determining the title level of the text line corresponding to the optimized feature vector in the first target title feature set; The first candidate optimized feature vector is stored in the first title feature set to obtain a first target title feature set, and the title level of the text line corresponding to the optimized feature vector in the first target title feature set is determined, including: Determine a reference feature vector corresponding to the first target similarity, search for a first cluster set corresponding to the determined reference feature vector, and store the first candidate optimized feature vector in the searched first cluster set; Determine the title level of the text line corresponding to the optimized feature vector in each updated first cluster set; wherein the updated first cluster sets are merged to obtain a first target title feature set.
13. The method according to claim 12, characterized in that The step of clustering the optimized feature vectors in the first title feature set to obtain a plurality of first cluster sets includes: Obtaining the serial number features contained in the optimized feature vector in the first title feature set, and determining whether the serial number features are the same; When the sequence number features are different, clustering the optimized feature vectors in the first title feature set according to the sequence number features to obtain a plurality of first cluster sets; Alternatively, when the sequence number features are the same, the optimized feature vectors in the first title feature set are clustered according to the font size features to obtain multiple first cluster sets.
14. The method according to claim 12, characterized in that The determining of the reference feature vector corresponding to each of the first cluster sets includes: For any of the first cluster sets, extracting reference features corresponding to the first cluster set, and forming a reference feature vector from the reference features; The reference feature includes at least one of the following: a reference font application feature, a reference bold feature, a reference serial number feature, a reference font size feature, a reference document boundary feature, a reference page index feature, a reference line index feature, a reference text feature, and a reference color feature.
15. The method according to claim 12, characterized in that In the case where the benchmark feature is a benchmark font application feature, the extracting of the benchmark feature corresponding to the first cluster set includes: counting the number of first font applications corresponding to the optimized feature vectors whose font application features are a first value in the first cluster set; counting the number of second font applications corresponding to the optimized feature vectors whose font application features are a second value in the first cluster set; in the case where the number of first font applications is greater than the number of second font applications, setting the benchmark font application feature corresponding to the first cluster set to the first value; or, in the case where the number of first font applications is not greater than the number of second font applications, setting the benchmark font application feature corresponding to the first cluster set to the second value; Alternatively, in the case where the reference feature is a reference bold feature, the extracting the reference feature corresponding to the first cluster set includes: counting a first bold number corresponding to the optimized feature vector whose bold feature is a first value in the first cluster set; counting a second non-bold number corresponding to the optimized feature vector whose bold feature is a second value in the first cluster set; in the case where the first bold number is greater than the second non-bold number, setting the reference bold feature corresponding to the first cluster set to the first value; or, in the case where the first bold number is not greater than the second non-bold number, setting the reference bold feature corresponding to the first cluster set to the second value; Alternatively, in the case where the reference feature is a reference serial number feature, the extracting the reference feature corresponding to the first cluster set includes: obtaining a serial number feature included in the optimized feature vector in the first cluster set, and setting the reference serial number feature corresponding to the first cluster set as the serial number feature; Alternatively, in the case where the reference feature is a reference font size feature, the extracting the reference feature corresponding to the first cluster set includes: obtaining the font size feature included in the optimized feature vector in the first cluster set, and determining an average font size feature of the font size features; setting the reference font size feature corresponding to the first cluster set as the average font size feature; Alternatively, in the case where the reference feature is a reference document boundary feature, the extracting the reference feature corresponding to the first cluster set includes: obtaining the document boundary feature included in the optimized feature vector in the first cluster set, and determining an average document boundary feature of the document boundary feature; setting the reference document boundary feature corresponding to the first cluster set as the average document boundary feature; Alternatively, in the case where the reference feature is a reference page index feature or a reference row index feature, the extracting the reference feature corresponding to the first cluster set includes: obtaining the page index feature included in the optimized feature vector in the first cluster set, and selecting the smallest page index feature from the page index features; setting the reference page index feature corresponding to the first cluster set as the smallest page index feature; obtaining the row index feature included in the optimized feature vector in the first cluster set, and selecting the smallest row index feature from the row index features; setting the reference row index feature corresponding to the first cluster set as the smallest row index feature; Alternatively, in the case where the reference feature is a reference text feature, the extracting the reference feature corresponding to the first cluster set includes: setting the reference text feature corresponding to the first cluster set to empty, where empty means that nothing is filled in; Alternatively, in the case where the benchmark feature is a benchmark color feature, extracting the benchmark feature corresponding to the first cluster set includes: obtaining the color feature contained in the optimized feature vector in the first cluster set, and determining the average color feature of the color feature; setting the benchmark color feature corresponding to the first cluster set as the average color feature.
16. The method according to claim 12, characterized in that The step of determining the title level of the text line corresponding to the optimized feature vector in each updated first cluster set includes: Sorting all the reference feature vectors to obtain a first sorting result; Determine, according to the first sorting result, the title level of the text line corresponding to the optimized feature vector in each updated first clustering set; The step of determining the title level of the text line corresponding to the optimized feature vector in each updated first clustering set according to the first sorting result includes: Traversing the reference feature vectors in the first sorting result in order; In the case where the traversed benchmark feature vector is the first benchmark feature vector, searching for the updated first cluster set corresponding to the traversed benchmark feature vector; determining the first-level title corresponding to the traversed benchmark feature vector as the title level of the text line corresponding to the optimized feature vector in the updated first cluster set corresponding to the traversed benchmark feature vector; Alternatively, when the traversed baseline feature vector is not the first baseline feature vector, the previous baseline feature vector of the traversed baseline feature vector is searched from the first sorting result; based on the previous baseline feature vector, the title level of the text line corresponding to the optimized feature vector in the updated first clustering set corresponding to the traversed baseline feature vector is determined.
17. The method according to claim 16, characterized in that The step of determining the title level of the text line corresponding to the optimized feature vector in the updated first clustering set corresponding to the traversed benchmark feature vector according to the previous benchmark feature vector includes: When the reference page index feature included in the previous reference feature vector is the same as the reference page index feature included in the traversed reference feature vector, determining the previous title level corresponding to the previous reference feature vector; The sum of the previous title level and the preset level threshold is determined as the title level corresponding to the traversed reference feature vector; Determine the title level corresponding to the traversed benchmark feature vector as the title level of the text line corresponding to the optimized feature vector in the updated first clustering set corresponding to the traversed benchmark feature vector; Add the traversed benchmark feature vector to the subtitle list of the previous benchmark feature vector; Alternatively, when the reference page index feature included in the previous reference feature vector is different from the reference page index feature included in the traversed reference feature vector, the optimized feature vectors in each updated first clustering set are sorted to obtain a second sorting result; For the updated first cluster set corresponding to the traversed reference feature vector, traverse the optimized feature vectors in the updated first cluster set, and search for the previous optimized feature vector of the traversed optimized feature vector from the second sorting result; In the case where the previous optimized feature vector does not exist, searching for the next optimized feature vector of the traversed optimized feature vector from the second sorting result; or, in the case where the next optimized feature vector exists, searching for the reference feature vector corresponding to the updated first clustering set where the next optimized feature vector is located; When the traversed baseline feature vector is different from the baseline feature vector corresponding to the updated first clustering set where the subsequent optimized feature vector is located, and the product between the baseline font size feature contained in the subsequent optimized feature vector and the preset first multiple is less than the baseline font size feature contained in the traversed baseline feature vector, determine the next title level of the baseline feature vector corresponding to the updated first clustering set where the subsequent optimized feature vector is located; The difference between the latter title level and the preset level threshold is determined as the title level corresponding to the traversed benchmark feature vector; the title level corresponding to the traversed benchmark feature vector is determined as the title level of the text line corresponding to the optimized feature vector in the updated first clustering set corresponding to the traversed benchmark feature vector; The reference feature vector corresponding to the updated first clustering set where the subsequent optimized feature vector is located is added to the subtitle list of the traversed reference feature vectors.
18. The method according to claim 17, characterized in that The method further comprises: In the case where there is a previous optimized feature vector, determining a previous title level of the reference feature vector corresponding to the updated first clustering set where the previous optimized feature vector is located; The sum of the previous title level and the preset level threshold is determined as the title level corresponding to the traversed reference feature vector; Determine the title level corresponding to the traversed benchmark feature vector as the title level of the text line corresponding to the optimized feature vector in the updated first clustering set corresponding to the traversed benchmark feature vector; The traversed base feature vector is added to the subtitle list of the next base feature vector.
19. The method according to claim 10, characterized in that The step of determining a second candidate optimized feature vector in the second candidate title feature set, storing the second title feature set in the second title feature set to obtain a second target title feature set, and determining a title level of a text line corresponding to the optimized feature vector in the second target title feature set includes: Clustering the optimized feature vectors in the second title feature set to obtain a plurality of second cluster sets; Traversing the optimized feature vectors in the second candidate title feature set, and determining a second similarity between the traversed optimized feature vector and a central optimized feature vector in any second cluster set; Finding the maximum second target similarity from the second similarities, and determining whether the maximum second target similarity is greater than a preset second similarity threshold; When the maximum second target similarity is greater than a preset second similarity threshold, determining the traversed optimized feature vector as a second candidate optimized feature vector; Storing the second candidate optimized feature vector in the second title feature set to obtain a second target title feature set, and determining the title level of the text line corresponding to the optimized feature vector in the second target title feature set; The step of storing the second candidate optimized feature vector into the second title feature set to obtain a second target title feature set, and determining the title level of the text line corresponding to the optimized feature vector in the second target title feature set includes: Finding a second cluster set corresponding to the largest second target similarity, and storing the second candidate optimized feature vector into the found second cluster set; The updated second clustering sets are merged to obtain a second target title feature set, and the title level of the text line corresponding to the optimized feature vector in the second target title feature set is determined.
20. The method according to claim 19, characterized in that The determining the title level of the text line corresponding to the optimized feature vector in the second target title feature set includes: Merging the first target title feature set with the second target title feature set to obtain a third target title feature set; Sorting the optimized feature vectors in the third target title feature set to obtain a third sorting result; Traversing the optimized feature vectors in the second target title feature set, and searching the position index corresponding to the traversed optimized feature vector in the third sorting result; According to the position index, the title level of the text line corresponding to the traversed optimized feature vector is determined.
21. The method according to claim 20, characterized in that Determining the title level of the text line corresponding to the traversed optimized feature vector according to the position index includes: Taking the position index as the center, forwardly search for the nearest neighbor forward optimization feature vector whose sequence number feature is the sequence number value in the third sorting result, and backwardly search for the nearest neighbor backward optimization feature vector whose sequence number feature is the sequence number value; According to the nearest neighbor forward optimization feature vector and the nearest neighbor backward optimization feature vector, the title level of the text line corresponding to the traversed optimization feature vector is determined.
22. The method according to claim 21, characterized in that The step of determining the title level of the text line corresponding to the traversed optimized feature vector according to the nearest neighbor forward optimized feature vector and the nearest neighbor backward optimized feature vector comprises: In the case where there is no nearest neighbor forward optimization feature vector but there is a nearest neighbor backward optimization feature vector, determining whether the traversed optimization feature vector is the optimization feature vector ranked first in the third sorting result; In the case where the traversed optimized feature vector is the optimized feature vector ranked first in the third sorting result, determine whether the first ratio between the font size feature in the traversed optimized feature vector and the font size feature in the nearest neighbor backward optimized feature vector is greater than the preset second font size threshold; in the case where the first ratio is greater than the preset second font size threshold, determine that the title level of the text line corresponding to the traversed optimized feature vector is a first-level title, and update the title level of the text line corresponding to the optimized feature vector whose sequence number feature is the sequence number value; or, in the case where the first ratio is not greater than the preset second font size threshold, determine the title level of the text line corresponding to the traversed optimized feature vector as the title level corresponding to the nearest neighbor backward optimized feature vector; or, When the traversed optimization feature vector is not the optimization feature vector ranked first in the third sorting result, search for the previous optimization feature vector of the traversed optimization feature vector in the third sorting result; determine whether the second ratio between the font size feature in the previous optimization feature vector and the font size feature in the traversed optimization feature vector is greater than the preset third font size threshold; when the second ratio is greater than the preset third font size threshold, determine the title level of the text line corresponding to the traversed optimization feature vector as the sum of the title level corresponding to the previous optimization feature vector and the preset level threshold; or, when the second ratio is not greater than the preset third font size threshold, determine the title level of the text line corresponding to the traversed optimization feature vector as the title level corresponding to the previous optimization feature vector.
23. The method according to claim 21, characterized in that The method further comprises: In the case where there is a nearest neighbor forward optimization feature vector but no nearest neighbor backward optimization feature vector, in the third sorting result, determining whether the nearest neighbor forward optimization feature vector is a previous optimization feature vector indexed by the position; In the case where the nearest neighbor forward optimization feature vector is the previous optimization feature vector indexed by the position, determine whether a third ratio between the font size feature in the traversed optimization feature vector and the font size feature in the nearest neighbor forward optimization feature vector is greater than a preset fourth font size threshold; in the case where the third ratio is greater than the preset fourth font size threshold, determine whether the text line corresponding to the traversed optimization feature vector is centered based on the document boundary feature in the traversed optimization feature vector; in the case where the text line corresponding to the traversed optimization feature vector is centered, determine that the title level of the text line corresponding to the traversed optimization feature vector is a first-level title, and update the title level of the text line corresponding to the optimization feature vector whose sequence number feature is the sequence number value; or, In the case where the nearest neighbor forward optimization feature vector is the previous optimization feature vector indexed by the position, determine whether the fourth ratio between the font size feature in the traversed optimization feature vector and the font size feature in the nearest neighbor forward optimization feature vector is less than the preset fifth font size threshold; in the case where the fourth ratio is less than the preset fifth font size threshold, determine the title level of the text line corresponding to the traversed optimization feature vector as the sum of the title level corresponding to the nearest neighbor forward optimization feature vector and the preset level threshold; or, in the case where the fourth ratio between the font size feature in the traversed optimization feature vector and the font size feature in the nearest neighbor forward optimization feature vector is not less than the preset fifth font size threshold, determine the title level of the text line corresponding to the traversed optimization feature vector as the title level corresponding to the nearest neighbor forward optimization feature vector; or, In the case where the nearest neighbor forward optimization feature vector is not the previous optimization feature vector of the position index, determine whether the fifth ratio between the font size feature in the traversed optimization feature vector and the font size feature in the previous optimization feature vector is less than the preset sixth font size threshold; in the case where the fifth ratio is less than the preset sixth font size threshold, determine the title level of the text line corresponding to the traversed optimization feature vector as the sum of the title level corresponding to the forward optimization feature vector and the preset level threshold; or, in the case where the fifth ratio is not less than the preset sixth font size threshold, determine the title level of the text line corresponding to the traversed optimization feature vector as the title level corresponding to the forward optimization feature vector; or, In the case where there is a nearest neighbor forward optimization feature vector and a nearest neighbor backward optimization feature vector, determine whether the reference feature vector corresponding to the nearest neighbor forward optimization feature vector is the same as the reference feature vector corresponding to the nearest neighbor backward optimization feature vector; if they are the same, determine the sixth ratio between the font size feature in the traversed optimization feature vector and the font size feature in the nearest neighbor forward optimization feature vector; determine the size between the sequence feature in the nearest neighbor forward optimization feature vector and the sequence feature in the nearest neighbor backward optimization feature vector; when the sixth ratio is greater than the preset seventh font size threshold and the nearest neighbor forward optimization feature vector When the serial number feature in the nearest neighbor backward optimization feature vector is smaller than the serial number feature in the nearest neighbor backward optimization feature vector, the title level of the text line corresponding to the traversed optimization feature vector is determined as a first-level title, and the serial number feature is updated to the title level of the text line corresponding to the optimization feature vector of the serial number value; or, when the sixth ratio is not greater than the preset seventh font size threshold, and / or the serial number feature in the nearest neighbor forward optimization feature vector is not smaller than the serial number feature in the nearest neighbor backward optimization feature vector, the title level of the text line corresponding to the traversed optimization feature vector is determined as the sum of the title level corresponding to the nearest neighbor forward optimization feature vector and the preset level threshold.
24. The method according to claim 23, characterized in that The method further comprises: In different situations, determine whether the title level corresponding to the nearest neighbor forward optimization feature vector is greater than the title level corresponding to the nearest neighbor backward optimization feature vector; in the case where the title level corresponding to the nearest neighbor forward optimization feature vector is greater than the title level corresponding to the nearest neighbor backward optimization feature vector, in the third sorting result, start with the nearest neighbor backward optimization feature vector, search backward for a new nearest neighbor backward optimization feature vector, and jump to the processing steps of the same situation, wherein the baseline feature vector of the new nearest neighbor backward optimization feature vector is the same as the baseline feature vector corresponding to the nearest neighbor forward optimization feature vector; Alternatively, when the title level corresponding to the nearest neighbor forward optimization feature vector is less than the title level corresponding to the nearest neighbor backward optimization feature vector, in the third sorting result, starting with the nearest neighbor forward optimization feature vector, a new nearest neighbor forward optimization feature vector is searched forward, and the process is jumped to the same situation, wherein the reference feature vector of the new nearest neighbor forward optimization feature vector is the same as the reference feature vector corresponding to the nearest neighbor backward optimization feature vector; Alternatively, when the title level corresponding to the nearest neighbor forward optimization feature vector is equal to the title level corresponding to the nearest neighbor backward optimization feature vector, the title level of the text line corresponding to the traversed optimization feature vector is determined as the sum of the title level corresponding to the nearest neighbor forward optimization feature vector and a preset level threshold.
25. A device for determining a document title level, characterized in that: The device comprises: A text line acquisition module is used to acquire a target document and a text line in a target document page of the target document, and store the text lines in a title candidate line set; A vector extraction module, used for extracting a basic feature vector and an optimized feature vector of each text line in the title candidate line set; A title determination module, used to determine the document title in the title candidate row set according to the basic feature vector and the optimized feature vector; The title level determination module is used to determine the title level of the document title.
26. A storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 24 is implemented.
Citation Information
Cited By
Multi-type text hierarchical directory construction method, device and equipment based on large model
CN121166840A