System and Method for Enabling Interlingual Information Access

TR202421327A2Pending Publication Date: 2026-07-21MUSAB GÜLTEKİN
0 Cites 0 Cited by

Patent Information

Application Number
TR202421327
Authority / Receiving Office
TR · TR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2026-07-21
Patent Text Reader

Abstract

The invention relates to a cross-language information retrieval system and method that optimizes language-specific processes, enables bilingual searching on mobile devices and personal computers across all operating systems, and specifically improves cross-language information retrieval between structurally different languages ​​such as Arabic and Turkish, and provides more accurate and comprehensive search results between different language pairs.
Need to check novelty before this filing date? Find Prior Art

Description

1 TARIFF Interlingual Information Access System and Method TECHNICAL AREA 5 The invention optimizes language-specific processes and works on mobile devices across all operating systems. Personal computers capable of bilingual searching, especially for languages ​​structurally similar to Arabic and Turkish. to improve cross-language knowledge access between different languages ​​and to foster greater understanding between different language pairs. 10. Enabling cross-language access to information that ensures accurate and comprehensive search results. It is related to the system and method. PREVIOUS TECHNIQUE Current bilingual search systems, especially those structurally compatible with languages ​​like Arabic and Turkish, are typically 15 languages ​​long. When it comes to different languages, it falls short in terms of accuracy and completeness. This Because systems often fail to adequately handle language-specific features, they miss relevant results. They tend to do so. This situation makes interlingual information access inefficient and incomplete, and is very problematic. This limits effective access to multilingual resources. Multilingual digital resources in the field of search and information access. Ensuring accurate and comprehensive access to cross-lingual information is critical for benefiting from resources. 20 This is important. However, the limitations of current bilingual search systems hinder this goal. This makes it difficult to accomplish. A major shortcoming of existing systems is their rigid, uniform approach. Systems often use general approaches that do not account for specific language pair complexities. 25 This situation leads to a lack of customized solutions tailored to language characteristics and different This leads to an inability to establish an accurate match between languages, especially those with a right-hand side like Arabic. In systems where left-handed writing and agglutinative languages ​​like Turkish coexist, this single type These approaches significantly reduce efficiency. Bilingual search systems often rely on the language... 30 Simple translation methods or general language models that ignore nuances and context This relies on the loss of semantic differences, which are critical for interlingual information access. Therefore, it leads to incorrect or incomplete results. For example, the word "çok" in Turkish... Meaningful words or root word structures in Arabic are accurate in simple translation methods. This may not be possible, especially when searching between structurally different languages. Another problem that leads to the loss of important information is the inadequacy of language-specific features. 35 It is the processing of the major differences between the linguistic structures of Turkish and Arabic, word order, etc. Additions and pronoun usage, for example, negatively affect the accuracy of search results. 2 This can have an impact. Therefore, systems that take into account the unique characteristics of each language are needed. It needs to be improved. Existing systems are generally not adaptable to the unique characteristics of different language pairs. Each Language structures, cultural differences, and modes of expression vary between language pairs, resulting in 5 existing examples. This prevents systems from working efficiently using generic templates. This situation allows users It is limited to content in only one language and cannot benefit from multilingual resources. This leads to a fundamental technical problem arising from these shortcomings: various language pairs. The problem is the inability to provide accurate and comprehensive cross-lingual information access between them. Search systems, To provide effective access to content in different languages, the structural features and context of the language are considered. It should be optimized to take into account language barriers and technical limitations. Otherwise, As a result, multilingual digital resources will remain inefficient, and information sharing between languages ​​will be limited. This situation will restrict access to information globally, hindering those who want to overcome language barriers. This poses a major challenge for users. Modern cross-lingual information retrieval solutions, It aims to overcome these problems and provide more effective and accurate access to information. 15 As a result of research conducted in the literature, the application number “2024 / 019321” and “DYNAMIC LANGUAGE” A Turkish patent application titled "TRANSLATION SYSTEM" has been found. The application allows users to automatically translate text content into their chosen language. It is related to the system. However, the application in question optimizes language-specific processes, and all operating systems... systems that can perform bilingual searches on mobile devices and personal computers, especially To improve cross-lingual knowledge access between structurally different languages ​​such as Arabic and Turkish. and obtaining more accurate and comprehensive search results across different language pairs. an indication of a system and method for providing interlingual information access No such encounters have been reported. 25 Ultimately, the problems mentioned above, which cannot be solved with current technology, are the subject of this technical analysis. This has made it necessary to make an innovation in the field. A BRIEF DESCRIPTION OF THE INVENTION The present invention aims to eliminate the aforementioned disadvantages and introduce new technologies to the relevant technical field. It relates to systems and methods of providing interlingual information access in order to bring advantages. The main goal of the invention is to optimize language-specific processes, making mobile applications available across all operating systems. enabling bilingual searching on devices and personal computers, especially between different languages ​​35 3 to improve cross-language information access and provide more accurate and comprehensive information between different language pairs. The goal is to provide search results. All the purposes mentioned above and those that will emerge from the detailed explanation below. The present invention optimizes language-specific processes to achieve this, and all operating systems are in operation 5. structured systems that can perform bilingual searches on mobile devices and personal computers. to improve cross-language knowledge access between different languages ​​and between different language pairs Cross-language information access that enables more accurate and comprehensive search results. It is a system for providing access to call requests made via an electronic device. Access via mobile device or personal computer in one of the supported languages ​​and UTF-8 10 a query input module that processes text input and cleans the input, detects the input language, and uses n-grams. Language detection uses analysis to examine character strings and calculate language probability. component that standardizes the input language, eliminates inconsistencies by applying normalization, and A query that prepares the search using language-specific character substitution rules. The normalization module is optimized according to the unique structures of the supported language, providing language-specific 15. Reverse indexing for efficient keyword searching, performing the initial search using algorithms. a primary search engine that uses and creates language-specific indexes, the primary search engine It classifies search results, distinguishing between core and compound entries according to linguistic rules. using pattern matching to identify compound inputs and categorize the results. The results categorization system prioritizes results from various search stages into 20 priority queues. a result that combines using, eliminates redundancies and consolidates the findings The merge and duplicate removal engine sorts the merged results by language pair and query. applying a custom comparator based on type-specific relevance criteria and the results The final result ranking algorithm uses an organized and efficient ranking algorithm. 25 It includes an electronic device interface. The best way to utilize the advantages of the existing invention, together with its structure and additional elements. For it to be understood, it must be considered together with the figures explained below. BRIEF DESCRIPTION OF THE FIGURES Figure 1 is a representative illustration of the cross-language information access system that is the subject of the invention. Figure 2 is a representative illustration of the method for enabling cross-language information access that is the subject of the invention. 35 The drawings do not necessarily need to be scaled and are necessary for understanding the invention. Details that are not present may have been overlooked. Furthermore, at least to a large extent... 4 Elements that are identical or at least have substantially identical functions are numbered the same. It is shown. REFERENCE NUMBERS 1. Query input module 2. Language perception component 3. Query normalization module 4. Primary search engine 5. Result categorization system 10 6. Advanced search module 7. Cross-language translation interface 8. Result merging and duplication removal engine. 9. Final result ranking algorithm 10. Electronic device interface 15 E. Yes H. No 1001. Search query from user via mobile hardware / electronic device 20 receiving and cleaning via the input module 1002. Detection of the input language's language perception component using n-gram analysis. to be done 1003. According to the detected language, the query is normalized via the query normalization module. normalization and standardization 25 1004. The primary search engine takes the normalized query and indexes it with language-specific indexes. performing the primary search 1005. Results categorization system performed by primary search engine. Categorizing search results using pattern matching 1006. If primary search results are insufficient, use the advanced search module 30. Implementing advanced search through 1007. Cross-language translation can be performed by the cross-language translation interface if needed. 1008. Combining and iterating over results from different search stages. Combining and eliminating duplicates via a debugging engine 1009. Final result ranking of combined and duplicate-free results 35 sorting / arranging according to relevance by algorithm 1010. Presenting the final results to the user via an electronic device interface. DETAILED DESCRIPTION OF THE INVENTION This detailed description outlines the system and method for enabling interlingual information access, which is the subject of the invention. 5 that will not create any limiting effect on a better understanding of the subject. This is explained with examples. The system for enabling interlingual information access consists of the query input module (1), language detection component (2), query normalization module (3), primary search engine (4), result categorization system (5), advanced search module (6), cross-language translation interface (7), result merging and iteration 10 removal engine (8), final result ranking algorithm (9), electronic device interface (10) It includes the query input module (1), search performed via electronic device. the request is made via a mobile device or computer in one of the supported languages ​​and It is the module that processes UTF-8 text input and cleans the input. Language detection component (2), input a language identifier that examines character strings using n-gram analysis and determines the language probability by 15 It is the unit that calculates. The query normalization module (3) standardizes the input language, Unicode eliminating inconsistencies by applying normalization and language-specific character modification. It is the module that prepares the search using the rules. Primary search engine (4), using language-specific algorithms optimized according to the unique structures of the supported language The first search is performed, reverse indexing is used for efficient keyword searching, and 20 are language-specific. It is the unit that creates the indexes. The result categorization system (5) is the primary search engine (4) It classifies search results, distinguishing between core and compound entries according to linguistic rules. using pattern matching to identify compound inputs and categorize the results. It is a system. Advanced search module (6), primary search engine (4) results are primary More sophisticated linguistic processes come into play when the results are insufficient. 25 implementing it, applying the Khoja root finder for Arabic and the deasciification process for Turkish. It is a module that performs more sophisticated linguistic operations (such as root extraction or deasciification). Cross-language translation interface (7), interface with translation services when monolingual searches are insufficient. It creates translations, makes API calls for translations, and caches frequently used translations. It is an interface that uses the mechanism. The result is the merging and redundancy removal engine (8), various 30 combines the results from the search stages using a priority queue, iterating through them. implements hash-based duplicate removal that eliminates and consolidates findings. It is a unit. The final result sorting algorithm (9) sorts the combined results, according to the language pair and It applies a custom comparator based on relevance criteria specific to the query type and presents the results. It is an algorithm that organizes and uses an efficient sorting algorithm. Electronic device 35 The interface (10) receives the sorted results from the final result sorting algorithm (9) 6 presents the results to the user, displays them, and allows the user to perform advanced search or translation. It is the device that allows it to be triggered. The system and method for providing cross-language information access is available on all operating systems such as Android and iOS. Multilingual digital mobile dictionaries that can perform bilingual searches on mobile devices in their systems, languages ​​5 multilingual research tools, multilingual e-commerce search engines, multilingual document management systems, cross-language patent search tools, cross-language social media monitoring tools, and much more. It is a system that can be used in various fields, such as multilingual legal document analysis applications. The system and method for providing interlingual information access addresses the problems mentioned in the previous technique. 10 It offers a gradual, language-specific approach to eliminating this. This approach involves searching. It adapts its methods to the unique characteristics of each language pair. This tailored (Specialized) process, especially in cross-language knowledge access between structurally different languages. It increases accuracy and comprehensiveness. It offers a system that enables cross-language information access. The elements, features, and algorithms used to provide the solution are as follows: 15  Query input module (1): The user's search request in one of the supported languages ​​on the mobile It is received via a device or personal computer.  Language detection component (2): Detects the input language, which is critical for subsequent processing. It has.  Query normalization module (3): Standardizes the input, eliminates inconsistencies and 20 They are ready to search.  Primary search engine (4): Optimized according to the unique structures of each supported language. It performs the initial search using algorithms specific to the developed language.  Result categorization system (5): Classifies search results, linguistically It distinguishes between kernel and compound inputs according to the rules. 25  Advanced search module (6): It activates if the primary results are insufficient, root extraction or It applies more sophisticated linguistic processes such as deasciification.  Cross-language translation interface (7): Translation services when monolingual searches are insufficient It creates an interface.  Result merging and duplication removal engine (8): 30 from various search stages It combines the incoming outputs, eliminates duplicates, and consolidates the findings.  Final result sorting algorithm (9): Combined results by language pair and query type It organizes according to specific relevance criteria.  Electronic device interface used for displaying results (10): Final, sorted It presents the results to the user. 35 7 The following processes can be carried out using the method of enabling interlingual information access:  The user enters the search query via a mobile hardware / electronic device. Receiving and cleaning via module (1) (1001),  Detection of the input language's language perception component (2) using n-gram analysis. (1002), 5  According to the detected language, the query is normalized via the query normalization module (3) normalization and standardization (1003),  The primary search engine (4) takes the normalized query and indexes it with language-specific indexes performing the primary search (1004),  The result categorization system (5) by the primary search engine (4) and 10 Categorizing the search results using pattern matching. (1005),  Advanced search module if primary search results are insufficient (6) Implementation of advanced search through (1006),  If necessary, cross-language translation can be done by the cross-language translation interface (7) 15 (1007),  Combining results from different search stages and removing duplicates. Combining and eliminating repetitions by means of the motor (8) (1008),  Final result ranking of combined and duplicate-free results sorting / arranging according to relevance by algorithm (9) (1009), 20  Presenting the final results to the user via the electronic device interface (10) (1010). The system for providing interlingual information access allows the user to enter their search query into the query input module (1) It receives it through this module. This module processes and cleans UTF-8 text input. Then, language detection... component (2) is activated and detects the input language using n-gram analysis. This process takes 25 This is critical for subsequent operations. According to the detected language, query normalization... module (3) standardizes the query. This module implements Unicode normalization and language-specific It uses character substitution rules. The normalized query is sent to the primary search engine (4) This engine uses reverse indexing for efficient keyword searching and language-specific indexes. It performs the first search by creating. Search results, result categorization system (5) 30 It is classified by this system. This system identifies compound entries using pattern matching and It categorizes the results. If the primary search results are insufficient, advanced search is used. module (6) is activated. This module activates the Khoja root finder for Arabic and for Turkish It applies more sophisticated linguistic operations such as deasciification. Monolingual searches are insufficient. When left unattended, the cross-language translation interface (7) is used. This interface makes API calls for translations 35 And it uses a caching mechanism for frequently used translations. From all these search stages... 8 The incoming results are processed by the result merging and redundancy removal engine (8). This The engine combines results using a priority queue and performs hash-based duplicate removal. It applies. Combined and duplicate-free results, final result ranking. It is regulated by the algorithm (9). This algorithm is a special algorithm based on relevance. It applies a comparator and uses an efficient sorting algorithm. Finally, the sorted 5 results are displayed via the user’s electronic device interface (10) is presented to the user. This interface also allows the user to perform advanced search or translation. It allows it to trigger its processes.

Claims

9 REQUESTS 1. Optimizing language-specific processes, on mobile devices across all operating systems, and Personal computers capable of bilingual searching, structurally distinguishing between different languages. To improve cross-language information access and to provide more accurate and 5-level information between different language pairs. Providing cross-language access to information that enables the acquisition of comprehensive search results. It is a system, and its feature is;  Supported call requests made via electronic device Accessible via mobile device or personal computer in one of the languages ​​and in UTF-8 Query input module (1) that processes text input and cleans the input, 10  detects the input language, examines character strings using n-gram analysis, and language perception component which calculates language probability (2),  Standardizes the input language, eliminates inconsistencies by applying normalization. and preparing it for search using language-specific character substitution rules. query normalization module (3), 15  Language-specific optimization based on the unique structures of the supported language. using algorithms to perform the initial search, for efficient keyword searching. The primary search engine that uses reverse indexing and creates language-specific indexes (4),  the primary search engine (4) classifies the search results, linguistically Pattern matching 20 distinguishes between kernel and compound inputs according to the rules. using results that identify composite inputs and categorize the results. categorization system (5),  Prioritizing results from various search stages using a queue a result that combines, eliminates redundancies and consolidates the findings Merge and redundancy engine (8), 25  Sorts combined results based on relevance specific to language pairs and query types. which applies a specific comparator according to its criteria and organizes the results and final result sorting algorithm using an efficient sorting algorithm (9),  the sorted results obtained from the final result sorting algorithm (9) electronic device interface (10) that presents to the user, 30 It includes.

2. It is a system that provides cross-language information access in accordance with Claim 1, and its feature is; as mentioned above. in case of insufficient primary results which are primary search engine (4) results more sophisticated linguistic processes such as root extraction or deasciification come into play. 35 It includes an advanced search module (6) which implements.

3. A system for providing interlingual information access in accordance with Claim 1, whose characteristic is that it is monolingual. When search results are insufficient, it interfaces with translation services and provides an API for translations. cross-platform that makes calls and uses a caching mechanism for frequently used translations. It includes a language translation interface (7). 5 4. Optimizing language-specific processes, on mobile devices across all operating systems, and Personal computers capable of bilingual searching, structurally distinguishing between different languages. to improve cross-language information access and to provide more accurate and reliable information between different language pairs. Providing cross-language access to information that enables comprehensive search results 10 It is a method, and its characteristic is;  The user enters the search query via a mobile hardware / electronic device. Obtaining and cleaning (1001) via the query input module (1),  Using n-gram analysis via the language perception component (2) of the input language detection (1002), 15  According to the detected language, the query is normalized via the query normalization module (3) normalization and standardization (1003),  The primary search engine (4) takes the normalized query and searches for language-specific Performing primary search with indexes (1004),  The result categorization system (5) by the primary search engine (4) 20 The search results were categorized using pattern matching. (1005),  Advanced search module if primary search results are insufficient. Implementation of advanced search via (6) (1006),  If needed, cross-language translation can be done by the cross-language translation interface (7) 25 to be done (1007),  Combining and iterating over results from different search stages Combining and eliminating duplicates by means of the elimination engine (8) (1008),  Final result ranking of combined and duplicate-free results 30 sorting / arranging according to relevance by algorithm (9) (1009),  Presenting the final results to the user via the electronic device interface (10) (1010) It includes the steps of the process. 35