DUPLICATE DOCUMENT GROUPING SYSTEM AND RELATED METHOD
Patent Information
- Application Number
- TR202421740
- Authority / Receiving Office
- TR · TR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2026-08-21
Smart Images

Figure 00000012_0000 
Figure 00000013_0000
Abstract
Description
1 TARIFF DUPLICATE DOCUMENT GROUPING SYSTEM AND RELATED METHOD Technical Area The invention uses artificial intelligence and machine learning techniques to recognize documents. It is focused on the processes of classification and interpretation. Specifically, document grouping 5. Analysis of visual and textual features within systems, document duplication the assessment of the situation and recording these documents in a database to facilitate operational processes a duplicate document grouping system that allows focusing on automation and It is related to the method. State of the Art 10 In the current known technology, manual classification and duplication of documents are common. Identifying documents is both a time-consuming and error-prone process. These challenges... Various methods and technologies have been developed to overcome this. Currently... In these techniques, methods that analyze the visual characteristics of documents are common. However, these The methods do not support text-based similarity analysis and are generally one-dimensional. It offers an approach. Optical Character Recognition (OCR) technology processes documents. It is widely used for digitization and text-based processing. However, The accuracy rate of OCR results remains low. In patent document number US10963691B2, this text describes how to document a device 20 It is explained that it classifies the data. The device obtains image data related to a document. Furthermore... Then, using this image data, it classifies the document into one of several document types. and determines a confidence score for this classification. It also provides a second It also determines a classification and a second confidence score related to it. The device uses first and second... It calculates the difference between the confidence scores. It compares this difference to a specific threshold value. 25 If the difference meets the threshold value, the device accepts the first classification. Patent document number US20220164370A1; for the classification of documents. It describes a method implemented by a computer. One or more processors By doing so, pre-trained classification models are collected in a model pool. The same In this way, various documents are collected in a document pool. These collected documents are then gathered by 30 pre-trained individuals. 2 Classification models are applied to documents in the document pool to create a tag list. is generated. Then, the final label result is processed by one or more processors. This is determined, and based on this final label result, a basis is established for document classification. An algorithm is created. The collected pre-trained classification models are documented. The application to the documents in the pool is carried out in parallel. Tag list 5 Estimation of text data retrieved from documents using a word length N. This is done by weighting the results. The final labeling result is determined by weighted voting. This is achieved by creating a strong label. Patent document number US20140136539A1; This text is classified as a visual document. It describes a system and method implemented by a computer. One or more 10 The resulting unencoded documents are each associated with a visual representation. Each with a classification code and a visual representation of that classification code. Related reference documents are also obtained. At least one of the unencoded documents, It is compared with reference documents, and based on this comparison, the unencoded document is created. Similar reference documents are identified. The classification codes of similar reference documents are 15. Based on this, a proposal is submitted to assign a classification code to the unencoded document. This proposal suggests that the visual representation of the proposed classification code should be presented alongside the unencoded document. It involves placing it within a portion of the associated visual representation. The proposed classification If the code is accepted, the visual representation of the accepted classification code. The size of the representation is increased. 20 As a result of the known state of the art, the aforementioned drawbacks occur. and due to the inadequacy of existing solutions on the subject, a need in the relevant technical field. Improvements were needed. Purpose and Brief Description of the Invention The purpose of the invention is to classify documents based on their visual and textual characteristics. 25 and enables the identification of duplicate documents. Another aim of the invention is to speed up business processes by eliminating manual document control. The goal is to minimize the time and human resources spent. Another aim of the invention is to integrate artificial intelligence and machine learning algorithms. This enables the accurate and rapid parsing of documents. 30 3 To achieve the above objectives, the present invention enables the detection of duplicate documents and documents to be classified and whether they are duplicates It is a document grouping system that includes a document database in which documents are recorded. - taking into account the visual and textual characteristics of the documents in the document database an identification module that is used for scanning and detection, 5 - data regarding document attributes obtained by the identification module Classification of documents by analyzing them using machine learning methods. a data processing module that provides, - Documents are checked for duplicates based on data processing module results. a grouping module that groups according to 10 It includes. The aforementioned identification module, - a visual algorithm service that extracts visual similarities between documents, - Levenshtein, Jaccard, to calculate the textual similarities of the documents. A 15 using the Levenshtein-Jaro and Levenshtein-Jaro-Winkler algorithms a text algorithm service that includes a service It includes. The aforementioned visual similarity algorithm layer includes SIFT and Image Hashes. The use of algorithms and Jaro-Winkler in the textual similarity algorithm layer. Jaccard similarity algorithms are applied. 20 Using the Random Forest algorithm by the machine learning module The documents are classified to determine if they are identical. 25 for grouping documents in a document database using artificial intelligence. method, - documents are recognized within an identification module and displayed on the user screen. Machine learning models to enable listing by name. using document-based separation, 4 - classified documents containing Optical Character Recognition (OCR) technology Original / copy differentiation is performed using a visual algorithm service, original and determining the number of copies and whether the documents are signed. determination, - Subsequently, using a visual algorithm service, the similarities between the documents were determined as 5. by comparing a group of documents, the unique key documents are identified. extracted and uploaded to a document recording application using OCR labeling, - a textual algorithm service that includes at least two different text similarity algorithms Calculating the intertextual similarity rate using, 10 - from scanned texts in documents via a data processing module or Questions and answers based on free text from sources such as instructions / Swift messages. preparing data sets by extracting information through this method and their labeling, - a grouping based on analysis results from the data processing module 15 Grouping documents according to their duplication status via the module. It includes the steps involved in the process. In addition, the obtained similarity data is processed through the data processing module. generating statistical values and applying the obtained values to an artificial intelligence model By using it in training, the system's accuracy rate is increased. Step 20 includes These advantages listed above both increase operational efficiency and... This also benefits the national and international banking sectors. Brief Description of the Figures Figure 1 shows the system components of the invention and the relationships between them. 25 It is seen. Figure 2 shows a flowchart illustrating the steps involved in the invention method. is provided. Reference Numbers 1. Document grouping system 30 10. Document database 20. Identification module 21. Visual algorithm service 22. Textual algorithm service 30. Data processing module 5 40. Grouping module 200. Document grouping method 201. Documents are recognized within an identification module and displayed on the user screen. using machine learning models to ensure they are listed by name document-based separation 10 202. Visual representations of classified documents containing Optical Character Recognition (OCR) technology. Original / copy distinction is made using an algorithm service, original and copy determination of the numbers and whether the documents are signed 203. By comparing the similarities of documents using a visual algorithm service. Extracting unique key documents from a group of documents and OCR 15 uploaded to and tagged in a document registration application using 204. A textual algorithm service that includes at least two different text similarity algorithms. calculating the intertextual similarity rate using 205. Scanned text from documents via a data processing module or Questions and answers based on free text from sources such as instructions / swift messages. 20 preparing and labeling datasets by extracting information through this method 206. A grouping based on the analysis results from the data processing module. Grouping documents according to their duplication status via the module. 207. The obtained similarity data is processed through the data processing module. generating statistical values and applying the obtained values to the artificial intelligence model 25 by using it in training to increase the accuracy rate of the system 6 Detailed Description of the Invention The system (1) components and the relationship between them relating to the system (1) which is the subject of the invention are shown in Figure 1. It is seen. The system in question (1) is based on the visual and textual features of documents. A document grouping system that enables the classification and identification of duplicate documents. 5 It is a system. The system for identifying and classifying duplicate documents (1) is generally: - a document in which documents to be determined whether they are duplicates are recorded document database (10) - visual and textual features of documents in the document database (10) 10 an identification module (20) which is scanned and detected taking into consideration - document properties obtained by the identification module (20) by analyzing the data using machine learning methods a data processing module that enables classification (30), - Based on the results of the data processing module (30), the documents are duplicated 15 a grouping module that groups according to its status (40) It includes. The aforementioned document database (10) is the starting point of the system. document database (10), It serves as a source for transferring documents to the document identification module (20). module (20) analyzes the documents in the database and processes them for other modules. It provides the data. The aforementioned identification module (20) visualizes the documents it receives from the document database (10). and analyzes its textual features. This module is divided into two sub-components: Visual Algorithm Service (21): SIFT and for analyzing the visual features of documents It uses Image Hashes algorithms. 25 Textual Algorithm Service (22): Calculating text-based similarities of documents Levenshtein, Jaccard, Levenshtein-Jaro and Levenshtein-Jaro-Winkler It includes algorithms. 7 The mentioned data processing module (30) analyzes the data coming from the identification module (20). It processes and assesses the possibility of duplication between documents. Data processing module (30) uses the random forest algorithm to determine whether the documents are the same. It classifies and processes data in a way that supports business processes. The mentioned grouping module (40) adds 5 to the results from the data processing module (30). Based on this, it groups documents according to their duplication status. Grouping module (40) improving the accuracy of the artificial intelligence model by evaluating statistical data It aims to... The system of detecting and classifying duplicate documents mentioned above (1) working method (200); 10 - documents are recognized within an identification module (20) and displayed on the user screen Machine learning models to enable listing by name. using document-based separation (201), - classified documents containing Optical Character Recognition (OCR) technology Original / copy distinction is made using the visual algorithm service (21), 15 Identifying the number of originals and copies, and verifying that the documents are signed. determining that they are not (202), - subsequently, the similarities of the documents were determined using the visual algorithm service (21). by comparing a group of documents, the unique key documents are identified. extracted and uploaded to a document recording application using OCR and 20 labeling (203), - a textual algorithm service that includes at least two different text similarity algorithms Calculation of intertextual similarity ratio using (22) (204), - from scanned texts in documents via a data processing module (30) or Questions and answers based on free text from sources such as instructions / swift messages 25 preparing data sets by extracting information through this method and their labeling (205), - a grouping based on the analysis results from the data processing module (30) through module (40) according to the duplication status of the documents grouping (206) 30 It includes the steps involved in the process. 8 In addition, the obtained similarity data are processed through the data processing module (30) generating statistical values and applying the obtained values to an artificial intelligence model (1) increasing the accuracy rate of the system (207) by using it in training It also includes that step. At the beginning of the method (200), documents are processed in an identification module (20). 5 The identification module (20) identifies document types using machine learning algorithms. It automatically identifies and displays the document types and names in the user interface. It automatically creates a list of document types such as commercial invoices, bills of lading, or certificates. It is distinguished as follows: To increase the accuracy of the machine learning model used. Training was conducted on a large and diverse dataset. 10 The identified documents were processed using a visual algorithm incorporating Optical Character Recognition technology. It is analyzed by service (22). This analysis performs the following operations: A distinction is made between original documents and copies. For example, a business invoice. Both the original and several copies can be submitted to the system; these documents are automated. It is distinguished as follows: 15 It is determined whether the documents are signed. This process ensures the legal validity of the documents. This is a critical step in this regard. This stage allows the system to perform specific processing for each document, and also... It reduces document duplication. Similarity analyses of documents are performed with the visual algorithm service (21). This analysis 20 As a result, a document group contains documents that are unique and not exactly the same as other documents. "Key documents" are selected. These documents are digitized using optical character recognition technology. It is transferred to the environment and saved by tagging it in a document management system (1). This step, It prevents unnecessary documents from being processed and ensures that critical documents are easily accessible. It makes it possible. 25 At this stage, a system containing at least two different algorithms is used to analyze the textual content of the documents. The textual algorithm service is started. The text similarity algorithms used are as follows: Levenshtein, Jaccard, 9 Jaro-Winkler. This method (200) allows for a deeper analysis of the contents of the documents. He recognizes. Through the data processing module (30), scanned texts from documents or free texts can be processed. Information is extracted from textual sources. This process, carried out using a question-and-answer method, involves 5 steps in the documentation. It ensures that critical information within is correctly labeled. The labeling process, It is necessary for training the machine learning model. Based on the analysis results obtained, the documents are grouped through the grouping module (40). They are grouped according to the degree of duplication. This both reduces the operational burden and It optimizes document management processes. 10 Finally, data such as similarity rates obtained from the documents are analyzed, and this The analysis results are used in training an artificial intelligence model. At the end of the training process, By increasing the accuracy rate of the model, documents can be accurately identified. Separation and classification are ensured.
Claims
REQUESTS 1. For the purpose of identifying and classifying duplicate documents, those that are duplicates a document database (10) in which documents that will be found to be missing are recorded a document grouping system (1) and its feature is; 5 - visual and textual features of the documents in the document database (10) an identification module (20) which is scanned and detected taking into consideration - document properties obtained by the identification module (20) by analyzing the data using machine learning methods a data processing module that enables classification (30), 10 - document duplication based on the results of the data processing module (30) a grouping module that groups according to its status (40) It includes.
2. The document grouping system in accordance with Claim 1 is (1), and its feature is the aforementioned definition. module (20) 15 - a visual algorithm service that extracts visual similarities of documents (21), - Levenshtein, Jaccard, to calculate the textual similarities of the documents. A algorithm that uses the Levenshtein-Jaro and Levenshtein-Jaro-Winkler algorithms. a textual algorithm service containing a service (22) It includes. 20 3. The document grouping system is (1) according to any of the previous requests, and its feature is; SIFT and Image Hashes algorithms in the visual similarity algorithm service (21) use of textual similarity algorithm service (22) with Jaro-Winkler This is an application of Jaccard similarity algorithms.
4. The document grouping system is (1) according to any of the previous requests, and its feature is; 25 Random Forest algorithm by machine learning module (30) It is the process of classifying whether documents are the same or not, using this method.
5. Grouping of documents (10) in a document database by means of artificial intelligence The working method (200) of a document grouping system (1) is its feature; 11 - documents are recognized within an identification module (20) and displayed on the user screen Machine learning models to enable listing by name. using document-based separation (201), - classified documents containing Optical Character Recognition (OCR) technology Original / copy distinction is made using the visual algorithm service (21), 5 Identifying the number of originals and copies, and verifying that the documents are signed. determining that they are not (202), - subsequently, the similarities of the documents were determined using the visual algorithm service (21). by comparing a group of documents, the unique key documents are identified. extracted and uploaded to a document recording application using OCR and 10 labeling (203), - a textual algorithm service that includes at least two different text similarity algorithms Calculation of intertextual similarity ratio using (22) (204), - from scanned texts in documents via a data processing module (30) or Questions and answers based on free text from sources such as instructions / swift messages 15 preparing data sets by extracting information through this method and their labeling (205), - a grouping based on the analysis results from the data processing module (30) through module (40) according to the duplication status of the documents grouping (206) 20 It includes the steps of the process.
6. The document grouping method according to claim 5 is (200), and its feature is the similarity obtained. The statistical values are processed through the data processing module (30) of the data. the creation and use of the obtained values in training the artificial intelligence model (1) Increasing the accuracy rate of the document grouping system by using (207) 25 It includes the process step.