A system for near language fine-grained recognition
Patent Information
- Application Number
- CN202610945526.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-29
- Publication Date
- 2026-09-25
AI Technical Summary
(1)频域噪音淹没问题:基于 n-gram 统计或词嵌入的主流语言检测工具(如FastText lid.176)在处理马来语与印尼语时,由于大量同源高频词(共有词)在频率域对两种语言的统计贡献几乎相等,模型无法从中获取有效区分信号,导致系统性错判——实验中对马来语长文本书籍的误判率超过 60%
1.从根源突破频域噪音壁垒:对数频域降权公式在数学层面将同源共有高频词的识别贡献精确归零,彻底消除了传统统计方法中同源噪音对近亲语言识别的系统性干扰;在马来语/印尼语长文本书籍级验证中实现 100% 正确识别率,相较 FastText 基线的系统性误判实现质的突破。
Smart Images

Figure CN122821576A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image and text multimodal understanding technology that deeply integrates natural language processing and computer vision. More specifically, it relates to a fine-grained automatic recognition system for cognate language pairs (especially Malay ms and Indonesian id) in optical character recognition (OCR) front-end scenarios. This system is driven by a dual-core approach, using glyph connectivity features (visual modality) in the physical domain of the image and logarithmically weighted marker word scores (language modality) in the frequency domain of the text, supplemented by strong prior rules for format symbols. It constructs a multi-channel heterogeneous feature weighted fusion decision framework, which is suitable for automatic language classification systems of cross-border trade documents, Southeast Asian invoices, and multilingual official document scans. Background Technology
[0002] In scenarios such as cross-border e-commerce, international logistics, and digital government in Southeast Asia, the system must automatically classify the language used in scanned documents, invoices, and official documents to complete subsequent field extraction, business routing, and compliance review. However, some languages belong to the same cognate family in linguistic classification and are extremely similar in written appearance and vocabulary, posing a systemic challenge to existing recognition technologies.
[0003] Malay (ms) and Indonesian (id) are a typical cognate language pair: they both belong to the Malayo-Polynesian branch of the Austronesian language family, share the same Latin alphabet writing system, and have an overlap rate of over 80% in their core everyday vocabulary (such as dengan, telah, untuk, adalah, etc.), making them almost indistinguishable at a statistical level.
[0004] Existing speech recognition technologies suffer from the following four core defects: (1) Frequency domain noise submersion problem: When processing Malay and Indonesian, mainstream language detection tools based on n-gram statistics or word embedding (such as FastText lid.176) cannot obtain effective distinguishing signals from the two languages because a large number of cognate high-frequency words (shared words) contribute almost equally to the statistical analysis of the two languages in the frequency domain, resulting in systematic misjudgment. In the experiment, the misjudgment rate of long Malay text books exceeded 60%.
[0005] (2) The problem of sparsity of features in very short texts: In scenarios such as invoice screenshots with only a few OCR fields, the statistical features of the text modality are extremely sparse, and no pure text method can give a reliable judgment.
[0006] (3) Single-modal blind zone problem: Existing methods rely entirely on the semantic modality of text and ignore the geometric features of characters in the physical domain of the image (character aspect ratio, connected structure, etc.). When the image quality deteriorates or the text is incomplete, the recognition rate drops sharply.
[0007] (4) Insufficient utilization of strong prior signals: There are high-confidence format symbols in specific language scenarios (such as the Indonesian Rupiah symbol Rp). Existing methods have not established corresponding symbol-language rule channels, resulting in missed opportunities for deterministic judgment in extremely short text scenarios. Summary of the Invention
[0008] The technical problem to be solved by this invention is to address the shortcomings of existing technologies by providing a fine-grained recognition system for close-kin language, specifically achieving the following three inventive objectives: First, break through the frequency domain noise barrier: design a logarithmic frequency domain weighting formula to mathematically attenuate common words in the frequency domain to zero, so that the system can automatically focus on language-specific marker words with high confidence on one side, thereby eliminating the interference of common words on the discrimination from the root. Second, construct a dual-modal image and text perception system: introduce a character-based connected domain analysis channel (visual modality) based on the physical domain of the image, and form a dual-core driven architecture that is independent and complementary with the logarithmic frequency domain text feature channel (language modality) to achieve robust recognition in scenarios with arbitrary text length; Third, establish a multi-channel heterogeneous feature weighted fusion framework: normalize and weight the image physical features, logarithmic frequency domain text features, format symbol strong prior features and baseline statistical model output, and ensure the overall recognition accuracy in complex real-world scenarios through multi-path redundancy.
[0009] The fine-grained recognition system for close-kin language of the present invention includes: The input layer is used to acquire the original scanned image and input the original scanned image into the input source to obtain a dual-output that includes the original scanned image and a list of row-level bounding boxes containing the coordinates and text content of each line of text region; The feature extraction layer has four feature extraction channels. The four feature extraction channels extract features from the dual-output and perform mapping and normalization processing on the features to obtain the confidence scores of each language under each feature extraction channel. The weighted linear fusion decision layer is used to perform weighted linear fusion processing on all the confidence scores corresponding to each language to obtain the total confidence score corresponding to each language; the language corresponding to the highest total confidence score is taken as the language recognition result.
[0010] Preferably, the four feature extraction channels include channel A, channel B, channel C, and the FastText baseline channel.
[0011] Channel A is used to extract features of the physical domain glyphs in the original scanned image. Channel B is used to identify the marker words in the row-level bounding box list by logarithmic frequency domain weighting; Channel C is used to perform strong prior rule recognition of format symbols on the list of row-level bounding boxes; The FastText baseline channel is used to obtain the FastText baseline statistical model score based on n-gram statistics or word embedding.
[0012] Preferably, in channel A, the following is performed: The image region is cropped from the original scanned image; The cropped image region is converted to grayscale and then binarized using Otsu's adaptive thresholding algorithm; Extract all connected components in the binarized image region and obtain the size parameter of each connected component; Based on the size parameters, noise filtering is performed on the connected components to obtain effective connected components, and the aspect ratio of all effective connected components is calculated. Then, the mean of all aspect ratios is calculated. Based on the normalized range in which the mean is located, the language bias score corresponding to each language is output as the confidence score of channel A.
[0013] Preferably, the size parameters include width, height, and pixel area; The noise filtering specifically involves discarding the current connected component if the width is less than a set width threshold, the height is less than a set height threshold, or the pixel area is less than a set area threshold.
[0014] Preferably, the original scanned image is cropped based on the coordinates in the row-level Bbox list.
[0015] Preferably, in channel B, the following is performed: The text content is segmented into words based on a pre-constructed word frequency mapping table to obtain vocabulary; The frequency of each word in each language is statistically analyzed, and the maximum cross-language frequency of the word is calculated based on the frequency. Calculate the logarithmic frequency domain weighted score of each word relative to each language based on the maximum frequency and the frequency. The logarithmic frequency domain weighted scores corresponding to all words in each language are summed and normalized to obtain the confidence score of channel B.
[0016] Preferably, the logarithmic frequency domain weighting score is calculated using the following formula: W(w_i, l) = log( f(w_i, l) / ( f_max(w_i) + ε ) ); In the formula, W(w_i, l) is the logarithmic frequency domain weighted score; w_i represents the i-th word; l represents the language type; f(w_i, l) is the frequency; f_max(w_i) is the maximum value of the frequency f(w_i, l) of language l; ε = 1e-9.
[0017] Preferably, in channel C, the following is performed: The text content is scanned using regular expressions based on a pre-constructed symbol-language rule table. If a symbol rule in the text content matches a symbol rule in the symbol-language rule table, the language corresponding to the symbol rule in the symbol-language rule table is extracted, and the channel C confidence score of that language is assigned a full score.
[0018] Preferably, the total confidence level is calculated using the following formula: Total_Score(lang) = 0.2 × S_FT + 0.1 × S_A + 0.5 × S_B + 0.2 × S_C; In the formula, S_FT is the confidence score of the FastText baseline channel; S_A is the confidence score of channel A; S_B is the confidence score of channel B; and S_C is the confidence score of channel C.
[0019] Preferably, the system further includes an output layer for outputting the language code, total confidence score, and confidence score of each channel corresponding to the language recognition result.
[0020] Beneficial effects The advantages of this invention are: 1. Breaking through the frequency domain noise barrier at its root: The logarithmic frequency domain weighting formula precisely reduces the recognition contribution of common high-frequency words at the mathematical level, completely eliminating the systematic interference of common noise on the recognition of closely related languages in traditional statistical methods; achieving a 100% correct recognition rate in Malay / Indonesian long text book-level verification, achieving a qualitative breakthrough compared to the systematic misjudgment of the FastText baseline.
[0021] 2. Image-text dual-modal collaborative perception: The visual modality (channel A-shaped physical features) and the language modality (channel B frequency domain weighted features) extract complementary information from two independent perception dimensions. It can rely on text frequency domain features for judgment when the image quality is poor, and can also rely on image physical features to assist decision-making in extremely short text scenarios. The two modalities serve as redundant backups for each other.
[0022] 3. Extremely low computational overhead, suitable for large-scale production deployment: Channel A is based on OpenCV connected component analysis (milliseconds, zero GPU required); Channel B is based on hash lookup of offline pre-computed weight tables (microseconds); Channel C is based on regular expression matching (microseconds). The entire multi-channel system requires no neural network inference and can support high-concurrency real-time processing in a regular CPU environment.
[0023] 4. Strong robustness in extremely short text scenarios: Channel C's strong prior mechanism for format symbols provides a deterministic judgment in the scenario of invoice screenshots containing only a few fields, making up for the inherent defect of confidence collapse of large language models in extremely short texts.
[0024] 5. High interpretability and auditability: Each judgment outputs an independent intermediate score for each channel, making the judgment process completely transparent and meeting the regulatory requirements of scenarios with high audit requirements such as finance and cross-border compliance.
[0025] 6. High scalability: The word frequency weight table of channel B supports incremental corpus updates; the rule table of channel C supports the addition of language symbols with zero code; the overall framework can be extended in parallel to any cognate close-relative language pair (such as written Persian and Dari). Attached Figure Description
[0026] Figure 1 This is a schematic diagram of the overall architecture of the close-kind language fine-grained recognition system of the present invention.
[0027] Figure 2 This is a schematic diagram of the feature extraction and mapping normalization process for channel A in this invention.
[0028] Figure 3 This is a schematic diagram of the feature extraction and mapping normalization process for channel B in this invention.
[0029] Figure 4 This is a comparison chart of the verification results of the system based on the present invention and the traditional FastText baseline for the recognition of three languages. Detailed Implementation
[0030] The present invention will be further described below with reference to embodiments, but this does not constitute any limitation on the present invention. Any limited modifications made by any person within the scope of the claims of the present invention are still within the scope of the claims of the present invention. The present invention discloses a fine-grained recognition system for close-kind language, which consists of four parallel feature extraction channels and a weighted fusion decision layer. The overall system architecture is as follows: Figure 1 As shown.
[0031] In this embodiment, the system's input source is the dual output of the OCR engine (PaddleOCR-VL, with mergeLayoutBlocks set to False): ① the original scanned image; ② a list of row-level bounding boxes (Bboxes), which contains the coordinates and text content of each line of text.
[0032] Based on the aforementioned dual-output, the system initiates four feature extraction channels in parallel. Each channel normalizes the feature mapping to the language confidence score within the [0, 1] interval. Finally, the recognition result is output by a weighted linear fusion decision layer. Total_Score(lang) = 0.2 × S_FT + 0.1 × S_A + 0.5 × S_B + 0.2 × S_C.
[0033] Wherein: S_FT is the FastText baseline statistical model score; S_A is the image physical domain glyph feature score (channel A, visual modality); S_B is the logarithmic frequency domain weighted marker word score (channel B, language modality core); S_C is the format symbol strong prior score (channel C). The sum of the weights of each channel is 1.0, with channel B having the largest weight (0.5), reflecting its position as the core of frequency domain precision analysis.
[0034] For the four feature extraction channels, namely channel A, channel B, channel C and FastText baseline channel.
[0035] Among them, the FastText baseline channel incorporates the existing open-source language recognition tool (FastText official pre-trained model lid.176.bin) as an independent statistical baseline channel into the multi-channel fusion framework. The purpose is to provide a cross-validation signal that is independent of the three innovative channels A, B, and C of this invention and has different error sources during the fusion stage, so as to enhance the overall robustness of the system.
[0036] The FastText baseline channel shares the same OCR output text with channels B and C, which is the text content corresponding to the bounding boxes output by PaddleOCR-VL. The FastText baseline channel takes a plain text string, either concatenated from one line or all lines of the text content, as input to the channel's FastText model's predict interface, allowing the model to return multiple candidate languages and their corresponding probability scores at once.
[0037] Channel A is used for extracting glyph features (visual modality) from the physical domain of the image.
[0038] Channel A extracts the geometric connectivity features of characters from the physical domain of the image. Its core finding is that there is a significant statistical separation between the average aspect ratio (W / H ≈ 0.9) of Austronesian agglutinative languages (Malay / Indonesian) and that of Indo-European analytic languages (English, W / H ≈ 1.2), which can serve as a low-computational-efficiency physical feature for distinguishing language families. Figure 2 As shown, the processing steps in channel A are as follows: First, the OCR engine outputs a list of row-level bounding boxes in "mergeLayoutBlocks: False" mode, and then crops out single-line image regions from the original scanned image based on the coordinates in the list of row-level bounding boxes.
[0039] Second, the cropped image area is converted to grayscale and then binarized using the Otsu adaptive thresholding algorithm to eliminate interference from different scanning brightness.
[0040] Third, call cv2.connectedComponentsWithStats (8-connected) to extract all connected components in the binarized image region, and obtain the position (x, y), width W, height H, and pixel area of each connected component.
[0041] Fourth, filter noisy connected components (discard connected components with W<5, H<5, or area<10), calculate the aspect ratio W / H for the remaining valid connected components, and take the average avg_ratio of all rows. For example, through actual processing, we obtained: Malay text avg_ratio ≈ 0.90, Indonesian ≈ 0.91, and English ≈ 1.22, showing significant differentiation among the three languages in this dimension.
[0042] Fifth, output the language bias score based on avg_ratio: avg_ratio>1.1 → strong bias towards English (S_A(en) = 0.9); avg_ratio ∈ [0.8, 1.0] → bias towards Austronesian languages (S_A(ms / id) = 0.6).
[0043] It should be noted that the connected component analysis in channel A can be replaced by a glyph embedding method based on deep learning, but the latter has a much higher computational cost than this method and is not suitable for production-grade lightweight deployment.
[0044] Channel B is used for identification of marker words through logarithmic frequency domain weighting (language modality core).
[0045] Channel B is the core innovation of this invention. Its essence is to reweight words in the frequency domain: through logarithmic transformation, the weight of "shared high-frequency words" is compressed to zero, while the weight of "specific high-frequency words" is exponentially amplified, achieving fine-grained disambiguation of closely related languages at the frequency domain level. For example... Figure 3 As shown, channel B specifically includes the following two stages.
[0046] In the offline pre-computation stage, based on the word frequency mapping table constructed from the word frequency statistics of representative corpora of each target language, the text content in the row-level Bbox list is segmented into words to obtain vocabulary. Then, the frequency f(w_i, l) of each word w_i in each language l is calculated, and the cross-language maximum frequency f_max(w_i) = max_l f(w_i, l) is calculated, that is, the maximum value f_max(w_i) is taken from the frequency f(w_i, l) for language l. The logarithmic frequency domain weighted score of word w_i for language l is calculated as follows: W(w_i, l) = log( f(w_i, l) / ( f_max(w_i) + ε ) ).
[0047] Where ε = 1e-9.
[0048] The frequency domain weighting mechanism of this formula can be analyzed as follows: When w_i is a high-frequency word shared across languages (e.g., dengan: high frequency in both Malay and Indonesian languages), f(w_i, l) ≈ f_max(w_i), then W → log(1) = 0, the weight of this word is automatically reduced to zero in the frequency domain, which is equivalent to eliminating the interference of shared words on recognition; When w_i is a high-frequency word specific to the target language (e.g., daripada is only high-frequency in Malay), f(w_i, ms)>>f(w_i, id), then W(w_i, ms)>>0. This word obtains a high distinguishing weight in the frequency domain and becomes a key signal for language discrimination.
[0049] During the online inference phase, after segmenting the OCR input text, it iterates through the pre-calculated weight table of the vocabulary query, accumulates the frequency domain weighted scores of vocabulary w_i for each language using the following formula, and normalizes the output confidence score S_B: S_B(l) = Normalize( Σ_imax(0, W(w_i, l)) ).
[0050] It should be noted that the logarithmic frequency domain weighting formula in channel B can be replaced by a TF-IDF variant or a mutual information (PMI) metric, but the logarithmic function has the advantages of low computational complexity and good smooth attenuation characteristics in mathematics, and experimental verification shows that it has the best convergence effect.
[0051] Experimental results show that introducing channel B completely eliminates FastText's misclassification of long Malay text books. The word "daripada" contributes 23.9 to the Malay frequency domain score, while the word "dengan" is weighted down to zero; ultimately, S_B(ms) = 0.92 and S_B(id) ≈ 0.01, achieving perfect disambiguation.
[0052] Channel C is used for strong prior rule recognition of format symbols.
[0053] Channel C utilizes highly specific format symbols (currency units, administrative codes, etc.) for specific language scenarios as strong prior judgment signals. In extremely short text scenarios (such as screenshots of invoices containing only a few fields), the confidence of Channel B is limited due to insufficient vocabulary samples, while Channel C can provide highly deterministic output.
[0054] The specific implementation of channel C is as follows: Maintain a symbol-language rule table, perform regular expression scanning on the input text content, and when a rule is matched, assign the corresponding language a score of S_C = 1.0 (full score); otherwise, assign 0 points. Example rule: Regular expression r'\bRp\.?\s [\d,.]+'Hit→ Strongly judge Indonesian (id), S_C(id) = 1.0; Regular expression r'\bRM\s [\d,.]+'Hit→ Strongly judge Malay (ms), S_C(ms) = 1.0.
[0055] The rule table supports dynamic loading at runtime, allowing for rapid expansion of rules for new languages or symbols without modifying the core algorithm.
[0056] In the weighted fusion decision layer of this system, the four-channel output scores are weighted and linearly fused using the Total_Score(lang) formula above, and the language with the highest total score is taken as the final language recognition result. The system also retains the intermediate scores of each channel for auditing and traceability. The final output structure includes: the predicted language code (ms / id / en), the total confidence of the fusion, the sub-scores of each channel, and the processing time.
[0057] Based on the above identification system, we selected some publicly available materials to verify the identification system.
[0058] Example 1: Precise disambiguation of closely related languages in long Malay text books Input: Nine pages were randomly selected from a Malay textbook (Bahasa Melayu), extracted using PaddleOCR-VL, and then entered into the recognition system. FastText baseline: It was misclassified as Indonesian by a slight difference of 0.491 (id) vs 0.429 (ms).
[0059] Channel A processing: Grayscale conversion, Otsu binarization, and cv2.connectedComponentsWithStats connected component extraction are performed sequentially on all line-level bounding box image regions in the 9 pages of the book. After filtering noise, the aspect ratio of all effective connected components is calculated, resulting in an avg_ratio ≈ 0.90 for the current input, falling within the Austronesian bias interval [0.8, 1.0]. Since printed Malay and Indonesian are highly similar in terms of glyph aspect ratio, Channel A can only distinguish between the Austronesian and English language families, and cannot further subdivide between Malay and Indonesian. Therefore, similar moderate bias scores are given to both: S_A(ms) = 0.45, S_A(id) = 0.43, S_A(en) = 0.05 (en is significantly suppressed due to the mismatch in aspect ratio features).
[0060] Channel B frequency domain weighting: For the word 'daripada' (a high-frequency word specific to Malay), f(daripada, ms) >> f(daripada, id), W(daripada, ms) = +23.9, contributing to a high score for Malay; for the word 'dengan' (a high-frequency word shared by Malay and Indonesian), f ≈ f_max, W → 0, automatic filtering, neither side's score increases. Output: S_B(ms) = 0.92, S_B(id) = 0.01.
[0061] Channel C processing: The concatenated text is scanned using regular expressions. If no currency symbols or format rules are found (the main text of the book does not contain exclusive symbols such as Rp and RM), Channel C maintains a neutral output: S_C(ms)=0.50, S_C(id)=0.50, S_C(en)=0.50, that is, the three languages are equally weighted, which does not produce biased interference to the fusion result. This conforms to the design principle of Channel C: "only gives a deterministic judgment when a strong prior rule is hit, and remains neutral when it is not hit".
[0062] Fusion decision (using Malay as an example): Total(ms) = 0.2×0.43 + 0.1×0.45 + 0.5×0.92 + 0.2×0.50 = 0.086 + 0.045 + 0.460 + 0.100 = 0.691; Total(id) = 0.2×0.49 + 0.1×0.43 + 0.5×0.01 + 0.2×0.50 = 0.098 + 0.043 + 0.005 + 0.100 = 0.246. The final correct classification is Malay (ms), disambiguation successful.
[0063] Example 2: Strong Prior Judgment of Format Symbols for Very Short Indonesian Invoices Input: Screenshot of Indonesian invoice. OCR output is "Harga: Rp. 10,000" (only 6 tokens). The text is too short, and neither FastText nor channel B can give a high confidence judgment.
[0064] Channel C processing: Regular expression r'\bRp\.?\s The string "Rp. 10,000" was matched by [\d,.]+', triggering the Indonesian Rupiah currency symbol rule and forcibly assigning S_C(id) = 1.0, S_C(ms) = S_C(en) = 0.
[0065] Fusion decision: Even though S_FT and S_B have low confidence levels due to text sparsity (approximately 0.4 each), the 0.2 × 1.0 = 0.20 points contributed by channel C have already created a significant score gap between the two languages, and the final correct judgment is Indonesian (id).
[0066] Example 3: Fast Image-Physical Domain Differentiation of English Text Input: Scanned copy of an English business invoice, which, after OCR extraction, contains a large number of English word image lines.
[0067] Channel A processing: After Otsu binarization of the English line image, connected components are extracted. English letters (such as W, M, H) generally have a large aspect ratio. The statistical avg_ratio = 1.28>1.1. The output of channel A is S_A(en) = 0.9 and S_A(ms / id) = 0.05.
[0068] Fusion Decision: The strong English physical signal from channel A, combined with the zero-match of Austronesian markers in the English text from channel B (S_B(ms) ≈ 0, S_B(id) ≈ 0), is determined to be English (en) with an absolute advantage. The total system response time is <50ms.
[0069] Example 4: System-level validation of nine-page sampling from three large books Validation scale: 9 pages were randomly selected from each of "Bahasa Melayu" (Malay), "Bahasa Indonesia" (Indonesian), and "English Grammar" (English), for a total of 27 pages. The validation was performed concurrently using validate_all_books.py.
[0070] like Figure 4 The figure shown is a comparison chart of the trilingual recognition verification results: a bar chart comparing the recognition accuracy of the traditional FastText baseline and the multi-channel fusion system of this invention in a 27-page book sampling test. Verification conclusion: All three languages were correctly recognized, with an accuracy rate of 27 / 27 = 100%. The FastText baseline method had a misclassification rate of approximately 55% for Malay / Indonesian. This method completely eliminates the systematic confusion problem of cognate and closely related languages through multi-channel fusion.
[0071] The above description is only a preferred embodiment of the present invention. It should be noted that those skilled in the art can make several modifications and improvements without departing from the structure of the present invention, and these will not affect the effectiveness of the implementation of the present invention or the practicality of the patent.
Claims
1. A fine-grained recognition system for close-kin language, characterized in that, include: The input layer is used to acquire the original scanned image and input the original scanned image into the input source to obtain a dual-output that includes the original scanned image and a list of row-level bounding boxes containing the coordinates and text content of each line of text region; The feature extraction layer has four feature extraction channels. The four feature extraction channels extract features from the dual-output and perform mapping and normalization processing on the features to obtain the confidence scores of each language under each feature extraction channel. The weighted linear fusion decision layer is used to perform weighted linear fusion processing on all the confidence scores corresponding to each language to obtain the total confidence score corresponding to each language; the language corresponding to the highest total confidence score is taken as the language recognition result.
2. The fine-grained recognition system for close-kin language according to claim 1, characterized in that, The four feature extraction channels include channel A, channel B, channel C, and the FastText baseline channel.
3. The fine-grained recognition system for close-kin language according to claim 2, characterized in that, In channel A, execute: The image region is cropped from the original scanned image; The cropped image region is converted to grayscale and then binarized using Otsu's adaptive thresholding algorithm; Extract all connected components in the binarized image region and obtain the size parameter of each connected component; Based on the size parameters, noise filtering is performed on the connected components to obtain effective connected components, and the aspect ratio of all effective connected components is calculated. Then, the mean of all aspect ratios is calculated. Based on the normalized range in which the mean is located, the language bias score corresponding to each language is output as the confidence score of channel A.
4. The fine-grained recognition system for close-kin language according to claim 3, characterized in that, The dimensional parameters include width, height, and pixel area; The noise filtering specifically involves discarding the current connected component if the width is less than a set width threshold, the height is less than a set height threshold, or the pixel area is less than a set area threshold.
5. A fine-grained recognition system for close-kin language according to claim 3, characterized in that, The original scanned image is cropped based on the coordinates in the row-level Bbox list.
6. A fine-grained recognition system for close-kin language according to claim 2, characterized in that, In channel B, execute: The text content is segmented into words based on a pre-constructed word frequency mapping table to obtain vocabulary; The frequency of each word in each language is statistically analyzed, and the maximum cross-language frequency of the word is calculated based on the frequency. Calculate the logarithmic frequency domain weighted score of each word relative to each language based on the maximum frequency and the frequency. The logarithmic frequency domain weighted scores corresponding to all words in each language are summed and normalized to obtain the confidence score of channel B.
7. A fine-grained recognition system for close-kin language according to claim 6, characterized in that, The logarithmic frequency domain weighting score is calculated using the following formula: W(w_i, l) = log( f(w_i, l) / ( f_max(w_i) + ε ) ); In the formula, W(w_i, l) is the logarithmic frequency domain weighted score; w_i represents the i-th word; l represents the language type; f(w_i, l) is the frequency; f_max(w_i) is the maximum value of the frequency f(w_i, l) of language l; ε = 1e-9.
8. A fine-grained recognition system for close-kin language according to claim 2, characterized in that, In channel C, the following is executed: The text content is scanned using regular expressions based on a pre-constructed symbol-language rule table. If a symbol rule in the text content matches a symbol rule in the symbol-language rule table, the language corresponding to the symbol rule in the symbol-language rule table is extracted, and the channel C confidence score of that language is assigned a full score.
9. A fine-grained recognition system for close-kin language according to claim 1, characterized in that, The total confidence level is calculated using the following formula: Total_Score(lang) = 0.2 × S_FT + 0.1 × S_A + 0.5 × S_B + 0.2 × S_C; In the formula, S_FT is the confidence score of the FastText baseline channel; S_A is the confidence score of channel A; S_B is the confidence score of channel B; and S_C is the confidence score of channel C.
10. A fine-grained recognition system for close-kin language according to claim 1, characterized in that, The system also includes an output layer, which outputs the language code, total confidence score, and confidence score of each channel corresponding to the language recognition result.