Method for automatically collecting and refining public legal data and a system utilizing the same

KR102999065B1Active Publication Date: 2026-08-03LAW COMPANY
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
LAW COMPANY
Filing Date
2026-03-10
Publication Date
2026-08-03

Smart Images

  • Figure 112026029043909-PAT00347_ABST
    Figure 112026029043909-PAT00347_ABST
Patent Text Reader

Abstract

A method for automatically collecting and refining public legal data according to one embodiment comprises the steps of: setting a source and document type to be collected; accessing at least one of a plurality of public legal data providers and requesting list information based on the set source and document type to be collected; selecting a new or updated target by comparing the list information with existing stored data; collecting detailed document information for the selected target; downloading a source file corresponding to the detailed document information or converting a non-standard document into text; generating structured document data from the detailed document information; generating integrated text based on the structured document data; supplementing the text by performing optical character recognition (OCR) on an image-containing area of ​​the integrated text; generating a source access URL and a search identifier based on the detailed document information, the structured document data, and the integrated text; verifying the validity of essential fields for the structured document data, the integrated text, the source access URL, and the search identifier to determine whether to save; and saving the structured document data and the source file, respectively, according to the result of the determination of whether to save. and includes the step of processing the stored legal data into an embedding and search index.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present invention relates to a method for the automatic collection and refinement of public legal data and a system utilizing the same. Background Technology

[0002] Public legal data, such as statutes, precedents, administrative rulings, authoritative interpretations, and local ordinances, is a core information asset widely utilized for determining citizens' rights and obligations, supporting legal practice, responding to regulations for businesses, and formulating intellectual property strategies. However, this data is fragmented among various providers, including national agencies, local governments, judicial and administrative bodies, and patent-related institutions; furthermore, the format of provision, search methods, document structure, and update cycles often vary among these providers. While some institutions offer public APIs, others provide information only through web page queries or file downloads, causing inconvenience for users who must individually search for and select the necessary materials.

[0003] Furthermore, even for identical or similar legal data, integrated utilization is difficult due to differences in document titles, case numbers, dates, body composition, attachment formats, encoding methods, and metadata items. Consequently, simple collection alone makes effective searching, analysis, comparison, and accumulation difficult, and refinement processes—such as removing duplicates, correcting missing values, parsing unstructured text, standardizing formats, and linking original texts with metadata—are essential. In particular, to continuously utilize large volumes of legal documents, it is becoming increasingly important to incorporate the latest data, exclude erroneous data, ensure consistency in storage structures, and process the data into a form that enables subsequent searching and analysis.

[0004] Therefore, the need for general crawling and data cleansing techniques to efficiently secure various public legal data and organize it into a usable format is continuously increasing. The problem to be solved

[0005] Public legal data is distributed across multiple agencies, including the National Law Information Center, the Intellectual Property Information Search Service, the National Tax Law Information System, and systems of individual ministries. Furthermore, data provision methods vary by agency—ranging from OpenAPI and webpage-based crawling to individual API integration—and response formats are diverse, including JSON, XML, HTML, HWP, and PDF, making it difficult to collect and process data uniformly. Additionally, document identification methods, metadata items, text structure, and the availability of original files differ by agency. Consequently, continuously securing necessary legal data requires performing customized collection, detailed information extraction, and text refinement procedures separately for each agency, resulting in a significant burden on development and operation. Moreover, simply collecting raw data is insufficient for immediate use in search or subsequent services; a refinement process is required to reconstruct scattered information—such as text, summaries, reasons, and footnotes—into integrated text, link and store original files with structured data, and filter out missing or inaccurate data. In particular, if the selection of new data through comparison with existing stored data, text extraction from source documents of different formats, validation of essential fields, and securing of a standardized storage structure are insufficient, problems such as duplicate collection, omitted collection, and accumulation of abnormal data may occur.

[0006] Therefore, the problem that the present invention aims to solve is to efficiently establish a foundation for the accumulation and utilization of public legal data by automatically collecting legal data of heterogeneous formats from multiple public legal data providers and structuring and refining it to store it in a consistent form.

[0007] The problems of the present invention are not limited to those mentioned above, and other unmentioned technical problems will be clearly understood by a person skilled in the art from the description below. means of solving the problem

[0008] A method for automatically collecting and refining public legal data according to an embodiment of the present invention for solving the above problem comprises, in a method for automatically collecting and refining public legal data performed on a server, the steps of: setting a source and document type to be collected; requesting list information by accessing at least one of a plurality of public legal data providers based on the set source and document type to be collected; selecting a new or updated target by comparing the list information and existing stored data; collecting detailed document information for the selected target; downloading a source file corresponding to the detailed document information or converting a non-standard document into text; generating structured document data from the detailed document information; generating integrated text based on the structured document data; supplementing the text by performing optical character recognition (OCR) on an image-containing area of ​​the integrated text; generating a source access URL and a search identifier based on the detailed document information, the structured document data, and the integrated text; and verifying the validity of essential fields for the structured document data, the integrated text, the source access URL, and the search identifier to determine whether to store them. The method includes the step of storing the structured document data and the original file, respectively, according to the result of determining whether to store them; and the step of processing the stored legal data into an embedding and search index.

[0009] The step of setting the source and document type to be collected above includes storing collection setting information for each of the plurality of public legal data providers, including an institution identifier, an institution name, key information used for generating integrated text, summary alternative field information, a document name extraction key, a document number extraction key, and setting information required for generating an original document access URL; and the step of accessing at least one of the plurality of public legal data providers based on the collection setting information and requesting list information, the step of collecting the detailed document information, the step of generating the structured document data, and the step of generating the original document access URL and search identifier. The step of selecting new or updated targets may include extracting a serial number, document ID, case number, or a composite identification value formed by combining two or more identification information included in the list information; comparing the extracted identification value with a corresponding identification value included in the existing stored data; determining a new document or an updated document as a selected target based on the comparison result; and collecting the detailed document information only for the selected targets.

[0010] The step of generating the integrated text comprises generating a header area including at least some of document name, institution name, document number, and date information from the structured document data; generating a body area including at least some of body text, summary, order, reason, background, appendix, and related laws from the structured document data; generating a footnote area including footnotes or additional explanations from the structured document data; and generating the integrated text by combining the header area, the body area, and the footnote area, wherein if a predetermined core field within the structured document data is empty, the integrated text is supplemented using the summary alternative field information; the step of generating the original text access URL and search identifier comprises generating a plurality of candidate original text access URLs and then determining the final original text access URL by comparing the title information or number information of the page indicated by the candidate original text access URL with the document information included in the detailed document information; the step of determining whether to save comprises determining whether to save by checking the existence and consistency of at least some of the document number, structured body data, integrated text, search identifier, and original text access URL; and the step of saving the structured document data and the original text file, respectively, comprises the structured Document data may be stored in a database, and the original file may be uploaded to an external storage and the storage path of the original file may be linked and stored in the database.

[0011] The step of selecting the new or updated targets above comprises each candidate document included in the list information above. Collector's value score regarding cast Calculating, and a pre-set processing budget The above collector value score within the range satisfying It includes selecting the new or updated target among the above candidate documents so that the sum of is maximized, and the above candidate documents is an individual public legal document identified from the above list information, and the above The above candidate document It is a recency score of 0 or more and 1 or less, calculated based on the time difference between the creation date, promulgation date, decision date, or update date of and the reference point, and the above is a novelty score of 0 or more and 1 or less calculated based on the result of comparison with the aforementioned existing stored data, and the above The above candidate document It is a source importance score of 0 or more and 1 or less, representing the importance of the collection target source to which belongs, and the above is an estimated utilization score of 0 or more and 1 or less calculated based on at least one of document type, organization type, historical search frequency, or reference frequency, and the above is a duplication risk score of 0 to 1 calculated based on at least one of the similarity of the title, number, body text, or date, and the above is a processing cost of 0 or more calculated based on at least one of API call volume, download volume, OCR execution volume, and parsing operation volume, and the above inside is a weight that adjusts the reflection ratio of each item, and the above processing budget may be a budget value set based on at least one of the maximum API call volume, maximum processing time, maximum storage capacity, and maximum computation volume.

[0012] The step of supplementing the above text is each document corresponding to the image-containing area. OCR reliability score regarding cast Calculated as, the above OCR reliability score The probability value that the OCR result is accurate based on Calculating, and the benefits obtained by reflecting the above OCR results and the loss resulting from misapplying the above OCR results Optimal threshold based on cast Calculating as, and the above The above It includes reflecting the above OCR result in the above integrated text or above body only when there is an abnormality, and the above is a value between 0 and 1, representing the word-unit average confidence of the OCR result, and the above is a value between 0 and 1, representing the degree of agreement between the OCR result and a pre-established legal terminology dictionary or legal context, and the above is a value between 0 and 1 indicating the degree of image noise, and the above is a value between 0 and 1 indicating the degree of distortion of character placement within the image or layout recognition error, and the above inside is a weight that adjusts the reflection ratio of each term, and the above is the above OCR reliability score It may be a probability value between 0 and 1, calculated by a monotonically increasing function that takes as input.

[0013] The step of generating the above original text access URL and search identifier is for each document corresponding to the above detailed document information. and multiple candidate URLs URL consistency score for each cast Calculated as, and the above URL consistency score It includes a step of confirming the candidate URL with the maximum value as the final source access URL, and said candidate URL is a candidate link for accessing the original text generated by applying multiple offset values ​​or multiple URL generation rules, and the above is the above candidate URL Title information of the page indicated by [the document] It is a value between 0 and 1 inclusive indicating the degree of agreement between the title information, and the above is the above candidate URL Page number information indicated by A and the above document It is a value between 0 and 1 inclusive representing the degree of agreement between the number information, and the above is the above candidate URL The institution information on the page indicated by A and the above document It is a value between 0 and 1 inclusive representing the degree of agreement between the institution information, and the above is the above candidate URL Date information of the page indicated by [the document] It is a value between 0 and 1 inclusive representing the degree of agreement between the date information, and the above is the above candidate URL is a penalty value between 0 and 1 indicating the degree to which it causes a redirect, error response, or abnormal path, and the above inside can be a weight that adjusts the reflection ratio of each term.

[0014] The step of processing into the above embedding and search index involves the chunk length of the stored legal data. and overlap length Each document divided according to Search relevance score for cast Calculated as, and the above search relevance score for the entire set of stored documents Optimal chunk length to maximize the average or sum and optimal overlap length Select and the optimal chunk length above and the optimal overlap length above It includes chunking the stored legal data based on [the data], generating an embedding vector, and loading it into a search index, and [the chunk length] is a value representing the number of characters or tokens included in each chunk, and the overlap length is a value representing the number of characters or tokens maintained to be duplicated between two adjacent chunks, and the above is the above chunk length and the above overlap length Documents under conditions It is a value between 0 and 1 inclusive representing the semantic cohesion within each chunk of, and the above is the above chunk length and the above overlap length It is a value between 0 and 1 inclusive indicating the degree to which the preceding and succeeding context is maintained under the condition, and the above is the above chunk length and the above overlap length It is a value between 0 and 1 inclusive that indicates the degree to which it contributes to search or query-response performance in the condition, and the above is the above chunk length and the above overlap length A value greater than or equal to 0 representing the indexing or search delay cost due to the condition, and the above is the above overlap length A value greater than or equal to 0 representing the cost of duplicate storage or duplicate calculation caused by, and the above inside can be a weight that adjusts the reflection ratio of each term.

[0015] A public legal data automatic collection and refinement system according to an embodiment of the present invention for solving the above problem comprises: a server including an AI model and a database; and a user terminal connected to communicate with the server. The server is configured to automatically collect public legal data from a plurality of public legal data providers, generate structured document data and integrated text for the collected public legal data, perform intelligent processing on the public legal data using the AI ​​model, and store the structured document data and the integrated text in the database. The user terminal is configured to transmit a request for inquiry, search, or management of public legal data to the server and output refinement results, search results, document body, or original text access information provided by the server.

[0016] The above AI model is configured to receive an image file for character recognition as input and output a recognized text string, or to receive chunk-unit legal text containing semantic context as input and output a high-dimensional embedding vector, and the server is configured to reflect the recognized text string in the integrated text or body and to generate a search index for the public legal data using the high-dimensional embedding vector.

[0017] The above AI model comprises as training data pairs of image files for character recognition extracted from legal documents and correct text strings corresponding to the image files for character recognition, and can be trained to update the parameters of the AI ​​model so that the loss value between the output of the AI ​​model, which receives the image files for character recognition as input and outputs a predicted text string, and the correct text string is reduced.

[0018] The above AI model can be trained to update the parameters of the AI ​​model according to a loss function that reduces the distance between embedding vectors for the positive learning pairs and increases the distance between embedding vectors for the negative learning pairs, by configuring chunk pairs that indicate the same document, the same event, or the same statute among chunk-unit legal texts containing semantic context as positive learning pairs and chunk pairs that indicate different documents, different events, or different statutes as negative learning pairs. Effects of the invention

[0019] According to an embodiment of the present invention, data can be automatically collected from multiple public legal data sources having different provision structures, thereby reducing the inefficiency of relying on manual searching and individual downloads by institution and enabling the continuous accumulation of a large amount of legal data. In addition, since only new or changed items can be selectively processed by comparing existing stored data with new data to be collected, duplicate collection can be suppressed and collection time and system resources can be reduced.

[0020] Furthermore, by reconstructing metadata, body text, footnotes, and related information obtained through detailed information retrieval into integrated text, and performing OCR processing on image areas or text conversion of unstructured document formats when necessary, it is possible to unify original texts of different formats into refined data that can be searched and subsequently utilized.

[0021] Furthermore, the reliability of stored data can be enhanced by filtering out abnormal data through the verification of essential fields such as number, body text, integrated text, search identifier, and original document URL. Additionally, by storing structured data in a database and original document files in an external repository, both accessibility to original documents and management convenience can be ensured. Ultimately, by integrating and automating the processes of collecting, refining, and storing public legal data, this invention has the effect of stably providing a data foundation suitable for subsequent applications, such as legal information services, search services, and analysis services.

[0022] The effects according to the embodiments are not limited to those exemplified above, and a wider variety of effects are included in this specification. Brief explanation of the drawing

[0023] FIG. 1 is a diagram illustrating the components of a public legal data automatic collection and refinement system according to one embodiment of the present invention. FIG. 2 is a diagram illustrating functional elements of a server according to one embodiment of the present invention. FIG. 3 is a diagram illustrating input and output values ​​through an AI model according to an embodiment of the present invention. FIG. 4 is a flowchart illustrating a method for automatically collecting and refining public legal data according to an embodiment of the present invention. FIG. 5 is a diagram illustrating the hardware configuration of a server according to one embodiment of the present invention. Specific details for implementing the invention

[0024] The advantages and features of the present invention and the methods for achieving them will become clear by referring to the embodiments described below in detail together with the accompanying drawings. However, the present invention is not limited to the embodiments disclosed below but may be implemented in various different forms. These embodiments are provided merely to ensure that the disclosure of the present invention is complete and to fully inform those skilled in the art of the scope of the invention, and the present invention is defined only by the scope of the claims.

[0025] Although terms such as "first," "second," etc., are used to describe various components, it goes without saying that these components are not limited by these terms. These terms are used merely to distinguish one component from another. Therefore, it goes without saying that the first component mentioned below may be the second component within the technical scope of the present invention. A singular expression includes a plural expression unless the context clearly indicates otherwise.

[0026] Hereinafter, embodiments of the present invention will be described in detail with reference to the attached drawings. Identical or similar reference numerals are used for identical components in the drawings.

[0027] FIG. 1 is a diagram illustrating the components of a public legal data automatic collection and refinement system according to one embodiment of the present invention.

[0028] A public legal data automatic collection and refinement system (hereinafter referred to as the system) may include an AI model (11), a server (10) including a database (20), and a user terminal (30).

[0029] The server (10) is a component that serves as the core control entity of the present invention and performs the functions of automatically collecting public legal data from external agencies, refining and structuring the collected data, and storing and providing the results. The server (10) can access multiple data sources to query list information, select new or updated targets by comparing them with previously stored data, and acquire detailed document information regarding the selected targets. Additionally, the server (10) can extract information such as title, case number, production date, summary, body text, and related laws from the acquired documents, convert them into structured data, and generate integrated text suitable for searching or subsequent use. Furthermore, the server (10) can verify the existence of essential fields to exclude abnormal data and control the storage of refined document data and original document access information in the database (20). Thus, the server (10) is not merely a simple communication relay device, but an application processing device that comprehensively performs functions such as data collection, duplicate determination, document parsing, text refinement, storage control, and user response provision.

[0030] As an example of implementation, the server (10) may be implemented as one or more physical servers, virtual machines, container-based cloud instances, or a combination thereof. For example, the server (10) may be implemented in a distributed manner as a collection server that performs crawling tasks, an application server that refines collected documents, a web server or API server that responds to user requests, and may also be implemented as a single integrated server as needed. Additionally, the server (10) may be configured to access public institution APIs or web pages at fixed intervals by including a scheduler or batch execution module, and may be implemented in a form that includes an OCR processing module, an embedding generation module, a PDF generation or upload module, a log management module, etc. In one embodiment, the server (10) may include a public institution OpenAPI linkage module, a web crawling module, and an individual service linkage module such as the Korean Intellectual Property Office, and may be implemented in a form that manages original files and structured data together by linking with a relational storage such as PostgreSQL and an external object storage.

[0031] The AI ​​model (11) is a component that performs intelligent processing functions to improve the refinement and usability of public legal data, which is called by the server (10) or executed within the server (10). The AI ​​model (11) can perform a character recognition function, for example, to convert character information contained in an image area within a legal document into text, and can perform a function to generate embedding vectors to express the refined legal document in semantic units. Accordingly, parts inserted in image form, parts of scanned documents, or unstructured expressions that are difficult to process with only plain text can be reduced to text-based data, and can support semantic-based search in subsequent search or question-and-answer systems. In other words, the AI ​​model (11) is a component responsible for converting the data collected by the present invention into a form that can be utilized by machines, rather than keeping it merely as a storage target.

[0032] In addition, the AI ​​model (11) is not necessarily limited to a single fixed algorithm, and different types of models may be selectively used depending on the purpose of application. For example, a model for character recognition, a model for legal document embedding, and a vector generation model reflecting document chunking results may be adopted. The AI ​​model (11) may be installed in the form of an internal library of the server (10), or it may be implemented by linking with an external AI API or a separate inference server. Furthermore, the AI ​​model (11) may be called depending on the type of document to be refined, the length of the document, the format of the information contained within the document, etc. For example, it may be configured to perform OCR inference only when an image exists, or to perform text embedding only when indexing for search is required. Therefore, the AI ​​model (11) functions as a selective or auxiliary intelligent processing means within the system of the present invention and is an element that improves refinement accuracy and utility.

[0033] The database (20) is a component that stores and manages public legal data collected and refined by the server (10). The database (20) may store structured data such as document titles, numbers, production dates, institution names, original URLs, document bodies, summaries, relevant laws, integrated text, and search identifiers. Additionally, the database (20) may store serial numbers, document IDs, case numbers, or composite identification values ​​to determine whether data has been newly collected, thereby serving as criteria for determining duplicates during subsequent re-collection. Furthermore, the database (20) may store additional data generated during the refinement process, such as management information including OCR results, blacklist status, original file storage paths, indexing targets, and embedding processing status. Thus, the database (20) is not merely a simple storage facility, but a storage means that maintains the history of automatic collection and refinement and forms a reference point for subsequent search, analysis, and provision.

[0034] In one embodiment, the database (20) may be implemented as a relational database management system and configured to store different types of legal data, such as rulings, interpretations, and adjudications, separately by table. Additionally, the database (20) may directly store the original file itself, but depending on the embodiment, PDF, HWP, or other original files may be stored in an external file storage or object storage, and the database (20) may be configured to store only the path or reference information of the file. Furthermore, the database (20) may be linked with a separate index storage for search services and may support updating or re-indexing existing data based on a unique document identifier when a document is modified or updated. Accordingly, the database (20) can be understood as a complex storage infrastructure for storing structured data, managing original document references, preventing duplication, and supporting subsequent searches.

[0035] The user terminal (30) is a component that communicates with the server (10) to perform requests for inquiry, search, viewing, or management of public legal data. The user terminal (30) allows the user to request desired legal information by inputting keywords, institution names, case numbers, date information, or document types, and can display on the screen the refined results, document body, original file links, search result lists, or analysis results returned from the server (10). Additionally, the user terminal (30) can provide an interface that allows the user to select and view only specific institutions, view detailed information and original files of specific documents, or download or rearrange search results based on user input. That is, the user terminal (30) is a device that forms a point of contact between the system and the user of the present invention, and is an element that provides a visual or interactive environment so that a person can verify and utilize the legal data processed by the server (10).

[0036] As an example of implementation, the user terminal (30) may be implemented as a desktop computer, a laptop computer, a tablet PC, a smartphone, a kiosk, or other network-connectable electronic device. Additionally, the user terminal (30) may operate as a web browser-based client program or may be implemented as a separate dedicated application. For example, a lawyer, patent attorney, corporate legal officer, researcher, or general user may access the server (10) through the user terminal (30) to check search results such as case law, judgment, authoritative interpretation, and local ordinance, and may use functions such as viewing the original PDF of a specific document, checking relevant laws and regulations, copying refined text, or linking to a subsequent Q&A service. Furthermore, the user terminal (30) may be implemented as a screen including multiple filter condition input windows, a search result list area, a detailed document viewer, and an original document download button, and is an interface device that performs request transmission and result reception by communicating bidirectionally with the server (10).

[0037] FIG. 2 is a diagram illustrating functional elements of a server according to one embodiment of the present invention.

[0038] Referring to FIG. 2, the server (10) may include a collection target management unit (201), a heterogeneous data source linkage collection unit (202), a list inquiry and collection target selection unit (203), a detailed document acquisition unit (204), a document parsing and refinement unit (205), an integrated text generation unit (206), an image character recognition processing unit (207), a source URL consistency verification unit (208), a validity determination and blacklist processing unit (209), a storage and file upload unit (210), a search identifier generation unit (211), and a post-processing embedding and indexing pipeline unit (212). Each component is not fixed as a single physical module but may be implemented as a program instruction, library, data structure, or distributed process, and as needed, two or more functions may be performed by a single module or a single function may be divided and performed by multiple modules.

[0039] The collection target management unit (201) is a functional element that manages standard information for collecting public legal data and institution-specific setting information. In one embodiment, the collection target management unit (201) can manage the identifier of the data provider institution, the institution name, the data type, key information to be used for assembling the text, the summary replacement field, the document name extraction key, the document number extraction key, and parameters required for generating the original text URL. For example, by holding setting values ​​such as target, institution_name, contents_text_keys, summary_keys, name_key, number_key, and url_data for a specific institution, subsequent modules can process different data within the same framework without hardcoding institution-specific differences. That is, the collection target management unit (201) functions as a meta-setting layer that absorbs structural differences in legal data for each institution and plays a role in increasing the universality and scalability of the entire server (10).

[0040] The heterogeneous data source linkage collection unit (202) is a functional element that accesses multiple different data providers to receive lists or detailed information. In the present invention, the data source is not limited to a single API, but may include a standardized public API such as the National Law Information Center OpenAPI, direct crawling of individual agency websites, and a specialized linkage interface such as the KIPRIS PLUS API. Accordingly, the heterogeneous data source linkage collection unit (202) can perform HTTP GET or POST-based API calls, HTML page requests, JSON / XML response reception, file download requests, etc., and can select an appropriate communication method according to the response format and access procedure for each source. Through this, the server (10) can collect legal data distributed across different agencies in an integrated manner without being dependent on a single source.

[0041] The list inquiry and collection target selection unit (203) is a functional element that retrieves list information from each data provider and selects the targets to be actually collected from among them. In one embodiment, the list inquiry and collection target selection unit (203) first retrieves the total count to calculate the number of pages, and then iterates through each page to extract the document serial number, document ID, case number, or other identification value. Subsequently, by comparing the identification value with the identification value in the previously stored database, only new documents that have not yet been stored or documents that require updating can be selected. Depending on the case, date filtering may be performed based on the range of the most recent modification date or publication date, or only specific categories may be selected as collection targets while iterating through categories by institution. In addition, a limit on the maximum number of collections per institution may be set to prevent a situation where the number of collected items increases excessively. In this way, the list inquiry and collection target selection unit (203) performs a core function that reduces duplicate collection and enables efficient incremental collection.

[0042] The detailed document acquisition unit (204) is a functional element that acquires detailed information for each selected document target. In one embodiment, the detailed document acquisition unit (204) can transmit a request to a detailed inquiry endpoint using a serial number, document ID, or event identification information obtained in the listing stage, and receive the full text of the document or detailed metadata from the response. At this time, the response format may be XML, JSON, HTML, or a file link, and may have different document structures such as resolutions, adjudications, decisions, interpretations, administrative appeals, or local ordinances depending on the characteristics of the institution. In addition, if original documents in PDF, HWP, HWPX, or DOC formats are provided, the corresponding file may be downloaded directly or a reference path may be secured for subsequent parsing. That is, the detailed document acquisition unit (204) is responsible for the function of acquiring original data and metadata that are the actual targets of refinement.

[0043] The document parsing and refinement unit (205) is a functional element that converts unstructured or semi-structured data collected by the detailed document acquisition unit (204) into a structured form. In one embodiment, the document parsing and refinement unit (205) can parse an XML response into a dictionary structure or extract items such as title, number, date, order, reason, background, appendix, and footnote from an HTML document using XPath or an HTML parser. Additionally, for non-standard files such as HWP, HWPX, and DOC, text can be extracted using a separate document parser or conversion tool. During the refinement process, elements that hinder subsequent searching, such as unnecessary spaces, special characters, residual tags, and null strings, can be removed or standardized. That is, the document parsing and refinement unit (205) performs the role of organizing data representations that differ by institution into a unified data structure.

[0044] The integrated text generation unit (206) is a functional element that generates integrated text suitable for search and question-and-answer based on refined structured fields. In one embodiment, the integrated text generation unit (206) can generate a single contents_text by combining three areas: a header, a body, and footnotes. For example, the header area may contain the agenda name, meeting type, organization name, etc., and the body area may sequentially combine key sections such as the summary of the decision, order, and reason. Additionally, if a specific key field is empty, alternative summary text can be generated using other fields corresponding to summary_keys, thereby mitigating information gaps caused by structural differences between documents. In this way, the integrated text generation unit (206) has the function of reconstructing the structure of the original document into a form that is easy for humans to read and easy for machines to search.

[0045] The image character recognition processing unit (207) is a functional element that restores character information existing in the form of an image within the document body or footnotes into text. In one embodiment, the image character recognition processing unit (207) is used in the integrated text generation process. It detects whether tags or equivalent image reference patterns exist and allows the corresponding image file to be downloaded separately. If the image format is unsuitable for direct OCR processing, such as GIF, it can be converted to another format, such as JPG, before performing OCR. The recognized text obtained as a result of OCR can be inserted into the body text location where the original image was situated or saved as separate OCR result data. Through this process, scanned images, text within figures, or image-based supplementary content can be converted into searchable text resources.

[0046] The source URL consistency verification unit (208) is a functional element that verifies whether the generated source access URL accurately points to the actual document. In one embodiment, the source URL consistency verification unit (208) assembles a basic URL pattern based on a serial number, but may sequentially apply multiple offset candidates to account for cases where the API-side serial number and the actual webpage URL index do not match. After actually accessing each candidate URL, the page title and document number are extracted and compared with the document name and document number already held by the server to determine consistency. If there is exactly one matching candidate, that URL is adopted as a valid source URL; if there is no match or multiple candidates match simultaneously, it can be processed as an abnormal or indeterminate state. Such a procedure corresponds to a verification function intended to ensure the accuracy of source access beyond simple URL combination.

[0047] The validity determination and blacklist processing unit (209) is a functional element that determines whether collected and refined document data is normal data that can be stored. In one embodiment, the validity determination and blacklist processing unit (209) can check for the existence of essential fields such as document number, structured contents, integrated text contents_text, search identifier, and original URL. If one or more of the essential fields are missing or invalid, the document can be determined as a blacklist target and excluded from normal storage targets. At this time, the blacklist determination history can be recorded as a log or status value and utilized for subsequent correction, re-collection, or management purposes. Accordingly, the validity determination and blacklist processing unit (209) performs the role of ensuring data quality and preventing incomplete documents from entering the service layer.

[0048] The storage and file upload unit (210) is a functional element that loads document data that has passed validation into a storage. In one embodiment, the storage and file upload unit (210) can store structured metadata, refined text, integrated text, date information, institution code, original document URL, blacklist status, etc., in a relational database such as PostgreSQL. Additionally, if an original document PDF or other attachment exists, it can be temporarily stored and then uploaded to an external object storage such as S3, and the path of the uploaded file can be linked and stored in a relevant field such as pdf_filepath in the database. Depending on the institution type or data format, HTML can be converted to PDF or downloaded original documents can be post-processed and uploaded. In this way, the storage and file upload unit (210) forms a dual storage structure of refined data and original documents, thereby simultaneously securing data usability and original document traceability.

[0049] The search identifier generation unit (211) is a functional element that generates an identification string to uniquely represent each document and improve search accuracy. In one embodiment, the search identifier generation unit (211) can generate a composite identification string, such as an expression, by combining multiple fields such as a document name, document number, serial number, and organization name. Unlike simple key values, this search identifier has the advantage of supporting both internal identification and external search simultaneously by reflecting document expressions that users are highly likely to actually search for. Additionally, the search identifier generation unit (211) can be utilized as an auxiliary key in the processes of duplicate document determination, log management, quality inspection, and search index creation.

[0050] The post-processing embedding and indexing pipeline (212) is a functional element that processes legal documents after collection and refinement are completed into a form suitable for semantic-based search. In one embodiment, the post-processing embedding and indexing pipeline (212) can extract legal data stored in a database and save it as a Parquet file, then divide it into chunks and generate an embedding vector for each chunk. In this process, an AI Embedding API or an equivalent embedding generation means may be utilized, and the generated chunk documents and vectors can be loaded first into Staging ES, then verified, and reflected in Production ES. Additionally, by caching intermediate results in a file system, re-calls can be reduced, and the same data can be maintained reproducibly. When a document is modified, existing chunks are deleted in bulk based on data_id and re-indexed with new chunks, thereby preventing the problem of old chunks remaining. That is, the post-processing embedding and indexing pipeline section (212) is responsible for the function of not merely storing the collection and refinement results of the present invention, but expanding them into search assets that can be utilized for semantic-based search and RAG-type legal services.

[0051] The server (10) can be understood not as a simple crawler, but as an integrated processing device that organically performs a series of functions, ranging from defining collection targets to accessing heterogeneous sources, selecting incremental collection targets, obtaining detailed original text, structuring and text refinement, OCR correction, checking URL consistency, determining data quality, storage and file management, generating search identifiers, and further embedding and search indexing. Through this structure, the server (10) can automatically accumulate public legal data distributed across multiple institutions and stably provide it in a form suitable for subsequent search, analysis, and response services.

[0052] FIG. 3 is a diagram illustrating input and output values ​​through an AI model according to an embodiment of the present invention.

[0053] Referring to FIG. 3, the AI ​​model (11) can be understood as being configured to receive an input value (410) and generate a corresponding output value (420). In this embodiment, the input value (410) may include at least a character recognition target image file (411) and a chunk-unit legal text (412) containing semantic context, and the output value (420) may include at least a recognized text string (421) and a high-dimensional embedding vector (422). That is, the AI ​​model (11) may function as an intelligent processing means that does not process only a single type of data, but converts different types of inputs into expressions suitable for their respective purposes.

[0054] The input value (410) refers to raw or intermediate processed data provided to the AI ​​model (11), which is data generated or extracted during the preprocessing or refinement process of the server (10). Among these, the image file (411) to be recognized as a character may be a file extracted from an image included in the body, footnotes, appendices, or other attachment areas of a legal document. For example, if an image tag is included within the legal document or if part of the scanned document exists in a non-textual form, the image may be used as input data for performing OCR. The image file (411) to be recognized as a character may be an image in GIF, JPG, PNG, or a similar format, and may be provided to the AI ​​model (11) after being converted into a format suitable for OCR as needed. This input is significant in that it serves as basic data for converting image-based character information, which must be read directly by a human, into text information that can be interpreted by a machine.

[0055] Additionally, the chunk-unit legal text (412) containing semantic context included in the input value (410) may be a text fragment obtained by dividing collected and refined public legal data into units of a certain length or semantic unit. Here, the chunk-unit legal text is not limited to a string of characters that merely cuts out a part of the text, but may be text that includes semantic context that aids in semantic interpretation, such as the title, name of the institution, case number, document type, context, and gist of the document to which the chunk belongs. Therefore, the chunk-unit legal text (412) containing semantic context can be understood as input data that mitigates the problem where the same legal term may be used with different meanings in different documents and enables the generation of semantic vectors that reflect the legal context. Such text input serves as a basis for generating expressions suitable for subsequent services such as search, question answering, and similar document search.

[0056] The AI ​​model (11) is an inference engine or intelligent processing module that receives the input value (410) and generates a corresponding output value (420). In this embodiment, the AI ​​model (11) is depicted as a single block in the drawing, but in actual implementation, it may be a single model, or a plurality of detailed models separated by purpose, or an external API call structure. For example, a character recognition model that processes a character recognition target image file (411) and an embedding generation model that processes chunk-unit legal text (412) may be implemented independently of each other, and logically, both may be included in the category of the AI ​​model (11). That is, the AI ​​model (11) of FIG. 3 may be a software module that runs on a single processor in terms of hardware, or it may be in a form that is linked to a cloud-based inference service or a separate AI server.

[0057] The output value (420) is result data that can be used in subsequent refinement, storage, indexing, or search processes of the server (10) as a result of processing by the AI ​​model (11). Among these, the recognized text string (421) may be string data generated as a result of reading characters included in the image file (411) subject to character recognition. For example, articles, headings, numbers, case numbers, ruling phrases, or footnote contents included within the image may be restored in the form of a text string. The generated recognized text string (421) may be inserted into the body text location where the original image was located, stored in a separate OCR result field, or passed to subsequent refinement logic. Through this, legal information existing in a non-text form can be converted into a text resource that can be searched and analyzed.

[0058] Meanwhile, the high-dimensional embedding vector (422) included in the output value (420) may be vector data that represents the meaning of a chunk-unit legal text (412) containing semantic context in a numerical space. This high-dimensional embedding vector (422) may be a dense vector composed of multiple dimension values ​​and may be generated to reflect the semantic similarity, contextual association, or query responsiveness of each text chunk. Accordingly, documents or sentences having substantially similar legal meanings, even if different expressions are used, can be placed adjacently in the vector space, and a subsequent search system can use this to perform a meaning-based search beyond keyword matching. Therefore, the high-dimensional embedding vector (422) is not merely data for storage, but a machine-friendly expression that can be utilized in indexing, similarity calculation, reordering of search results, and RAG-based question answering.

[0059] The relationship between the illustrated input value (410) and output value (420) illustrates an example where a character recognition target image file (411) corresponds to a recognized text string (421), and a chunk-unit legal text (412) containing semantic context corresponds to a high-dimensional embedding vector (422), but the present invention is not limited thereto. That is, two or more output values ​​may be generated from a single input value, or results of different formats may be generated even from the same type of input depending on the preprocessing state. Additionally, if necessary, the output value (420) may be post-processed by other components of the server (10) and stored in a database or loaded into a search index. Ultimately, the AI ​​model (11) of the present invention can be understood as a structure that improves the accuracy of legal data refinement and subsequent usability by converting unstructured input of public legal data into text-based results or semantic-based vector results.

[0060] In one embodiment, if the AI ​​model (11) is a character recognition model that generates a recognized text string (421) from a character recognition target image file (411), the AI ​​model (11) may be pre-trained using a training dataset that includes a pair of a character recognition target image file (411) extracted from a legal document and a correct text string corresponding to the character recognition target image file (411). During the training process, the AI ​​model (11) receives the character recognition target image file (411) as input, outputs a predicted text string, and parameters may be updated so that the loss value between the predicted text string and the correct text string is reduced. Accordingly, the AI ​​model (11) can more accurately restore character information contained in the image area within the legal document.

[0061] In another embodiment, where the AI ​​model (11) is an embedding generation model that generates high-dimensional embedding vectors (422) from chunk-unit legal texts (412) containing semantic context, the AI ​​model (11) may be trained using multiple chunk-unit legal texts. For example, pairs of chunk-unit legal texts that are included in the same document or refer to the same event or the same statute may be composed of positive learning pairs, and pairs of chunk-unit legal texts that refer to different documents, different events, or different statutes may be composed of negative learning pairs. The parameters of the AI ​​model (11) may be updated according to a loss function that decreases the distance between embedding vectors for the positive learning pairs and increases the distance between embedding vectors for the negative learning pairs. Accordingly, the AI ​​model (11) can generate high-dimensional embedding vectors (422) that more precisely reflect legal context and semantic similarity.

[0062] FIG. 4 is a flowchart illustrating a method for automatically collecting and refining public legal data according to an embodiment of the present invention.

[0063] Referring to FIG. 4, the method for automatically collecting and refining public legal data may include the steps of: setting a source and document type to be collected (S301); accessing at least one of a plurality of public legal data providers and requesting list information (S302); selecting new or updated targets by comparing with existing stored data (S303); collecting detailed document information for the selected targets (S304); downloading original files or converting non-standard documents into text (S305); generating structured document data from the detailed information (S306); generating integrated text based on the structured document data (S307); performing OCR on image-containing areas to supplement the text (S308); generating original access URLs and search identifiers (S309); verifying the validity of required fields and determining whether to save (S310); saving structured data and original files respectively (S311); and processing the saved legal data into embeddings and search indexes (S312). Through this series of steps, the present invention aims to automatically collect public legal data having different formats and structures depending on the institution, refine and structure it into a form that can be searched and subsequently utilized, and further process it into a data asset suitable for semantic-based search.

[0064] The step of setting the source and document type to be collected (S301) serves as the starting point of the entire collection pipeline. In this step, the server can determine which institution or system to collect data from, and what type of document to collect belongs to, such as legal precedents, authoritative interpretations, administrative appeals precedents, local ordinances, or patent-related legal data. Additionally, to handle data structures that differ by institution, the server can set the institution code, institution name, document name extraction key, document number extraction key, fields to be used for body composition, summary replacement fields, and additional parameters required for URL generation. In other words, step (S301) can be viewed as a pre-definition step that enables subsequent steps to operate within a common framework without individually reinterpreting differences between institutions. If this step is properly executed, it becomes possible to flexibly process the different document structures of different institutions using the same collection engine.

[0065] The step (S302) of accessing at least one of a plurality of public legal data providers to request list information is a step of initiating communication with an actual external data source. In this step, the server may access different providers, such as OpenAPI, individual agency websites, or specialized API services. The connection method may be an HTTP-based API request, an HTML page request and parsing, or, in some cases, a method of receiving a file download link. When requesting list information, parameters such as response format, sorting criteria, page size, agency code, authentication key, and date filter may be used together. In particular, if configured to first determine the total number of cases and then calculate the required number of pages, a large amount of data can be collected sequentially in a stable manner. Ultimately, the step (S302) is a step of securing list data that forms the scope of the actual collection target, and serves as a prerequisite for subsequent selection and detailed inquiry.

[0066] The step of selecting new or updated targets by comparing with existing stored data (S303) is a key step that enables incremental collection. In this step, the server can extract serial numbers, document IDs, case numbers, or composite identifiers included in the list information and compare them with identifiers previously stored in the database. As a result, only new documents that have not yet been stored or documents with the potential for updating can be selected as targets for subsequent collection. If necessary, the last modification date, publication date, category by institution, and a limit on the maximum number of collected items may also be considered. The technical significance of this step lies in reducing the inefficiency of re-collecting every document every time, saving system resources and processing time, while maintaining up-to-dateness. Furthermore, if a single serial number is insufficient depending on the institution, the possibility of duplicate or incorrect collection can be reduced by considering additional keys such as billing numbers or URL patterns.

[0067] The step of collecting detailed document information for selected targets (S304) is a step of expanding list-level summary information into actual usable document-level information. In this step, the server sends a request to a detailed inquiry endpoint or a detailed webpage for each selected target and can receive the full text of the document or detailed metadata from the response. For example, the case name, case number, resolution date, summary of decision, order, reason, appendix, relevant laws and regulations, and attached material information may be collected in this step. Depending on the institution, the response format may be XML, JSON, HTML, or a separate link structure, and the document may have an intermediate wrapper structure. Therefore, step (S304) can be described as a step of securing actual document data that serves as the basis for subsequent refinement by interpreting the detailed structure for each institution, rather than simply collecting links.

[0068] The step of downloading original files or converting non-standard documents into text (S305) involves securing the original text itself, in addition to detailed information, and converting it into a machine-processable format. Some institutions may provide only original text in file formats such as PDF, HWP, or HWPX, without providing direct body fields. In such cases, the server can download the file and store it in a temporary repository or secure it for subsequent uploading. Additionally, text can be extracted from non-standard or binary documents using a separate parser or conversion tool. This step takes into account that the actual sources of public legal data are not always provided solely as structured API responses; it serves to increase overall data coverage by including file-based documents within the scope of collection.

[0069] The step of generating structured document data from detailed information (S306) is a step of arranging semi-structured or unstructured information into a consistent internal schema. In this step, the server can generate structured data objects by extracting items such as title, number, date, institution name, body section, additional description, and relevant laws from the detailed document information. Since information with the same meaning may be provided with different field names depending on the institution, it is desirable to map it to internal fields by referring to the key information set in step (S301). Additionally, date format normalization, removal of unnecessary spaces, cleanup of special characters, and conversion to a JSON serializable form can also be performed in this step. Through step (S306), subsequent steps can operate based on a consistent data structure rather than raw responses specific to each institution.

[0070] The step of generating integrated text based on structured document data (S307) is a step of reorganizing structured fields into a single text resource suitable for searching and answering. In one embodiment, the integrated text may be generated by dividing it into header, body, and footnote areas. For example, the header may include document name, meeting type, organization name, etc., and the body may sequentially combine summary of decision, order, reason, background, appendix, etc. If a specific key field is empty, the gap may be filled by combining a replaceable summary field. This step is important in that it converts the field-dispersed structure of the original document into integrated text that is easy for both humans and machines to utilize. In other words, step (S307) is a step of reprocessing raw metadata in a way that enhances search suitability and contextual clarity, rather than merely storing it.

[0071] The step of supplementing the main text by performing OCR on image-containing areas (S308) is a correction step that converts non-text areas within the document into text resources. The main text or footnotes of legal documents may contain notations inserted in image form, scanned areas, text within figures, attached images, etc. In this case, the server can detect image reference tags or image areas, download the corresponding image files, and perform OCR after format conversion as necessary. The recognized text generated as a result can be inserted at the location where the original image was or saved as separate OCR result data. This step has the effect of enhancing the completeness of the integrated text and preventing subsequent search omissions by converting information that remained as images into searchable text. Therefore, although the step (S308) may be an optional supplementation step, it is a very useful step in terms of actual service quality.

[0072] The step of generating an original document access URL and a search identifier (S309) is a step that simultaneously ensures the traceability and searchability of the collected documents. First, the original document access URL is link information that allows a user to directly access the National Law Information System or the detailed page of the relevant agency. However, since the serial number on the API may not always match the actual webpage URL number, a procedure may be involved to assemble the URL by applying multiple offset candidates and verify whether the actual page title and number match the document information held. Meanwhile, the search identifier may be a composite string combining the document name, document number, serial number, agency name, etc., which supports both internal identification and external search simultaneously. Step (S309) is significant in that it is a step that considers both document consistency verification and search precision improvement, rather than simply generating a link.

[0073] The step of verifying the validity of essential fields and determining whether to save (S310) is a gateway step designed to ensure data quality. In this step, the server can check whether all fields essential for storage and service, such as document number, structured body data, integrated text, search identifier, and source URL, are validly secured. If one or more fields are missing or abnormal, the document may be classified as a blacklist target to be excluded from storage or a warning log may be generated. Conversely, if the necessary conditions are met, it may be classified as normal data and proceed to the subsequent storage procedure. This step serves to prevent incomplete data from entering the service layer and to preemptively reduce deterioration of search quality or user confusion.

[0074] The step of storing structured data and original document files (S311) separately is a step of fixing the refinement results as actual system assets. Structured data can be stored in a relational database, etc., and may contain various fields such as title, number, date, organization name, body text, integrated text, URL, blacklist status, OCR result, and file path. Meanwhile, original document files can be uploaded to an external object storage or file storage, and path information referencing the storage location can be stored together in the database. By adopting this dual storage structure, structured data suitable for searching and statistics can be secured on the one hand, while traceability that allows users to directly verify the original documents can also be ensured on the other. Therefore, step (S311) can be understood as a step where structured data management and original document evidence management are combined.

[0075] The step of processing stored legal data into embeddings and search indexes (S312) is an advanced post-processing step of the present invention. In this step, the server can divide the stored legal documents into chunks, add semantic context such as document title, institution name, and document type as needed, and then generate embedding vectors for each chunk. The generated vectors and chunk documents are loaded into a search index and can be utilized for semantic-based search, similar document search, and RAG configuration in a question-and-answer system. Additionally, if the data is updated, it is possible to delete the existing index and reload it, or to upsert based on document identifiers. Step (S312) demonstrates the scalability of the present method in that it does not end with collection and refinement, but converts the stored legal data into search assets that can be utilized in actual intelligent services.

[0076] Ultimately, each step of the method for the automatic collection and refinement of public legal data is organically connected through a flow of setup, list acquisition, incremental selection, detailed acquisition, file acquisition, structuring, integrated text conversion, OCR enhancement, URL and identifier generation, validation, storage, and embedding and indexing. Accordingly, this invention enables the implementation of public legal data—from collection to refinement, storage, and advanced search—as a single continuous processing system, thereby mitigating limitations in utilization caused by institutional format differences and unstructured nature. This sequential processing structure can be considered a methodology that substantially contributes to the mass accumulation of legal data, continuous updating, improved search quality, and the securing of a foundation for subsequent legal AI services.

[0077] Meanwhile, through certain steps of the automated collection and refinement method for public legal data, it is possible to generate a legal data set that maximizes recency, non-redundancy, document consistency, OCR accuracy, and searchability under limited collection and processing resources. Rather than simple automated collection, it is possible to mathematically determine "which documents to collect first, which documents to adopt as valid data after collection, which OCR results to reflect, and which chunking conditions are most advantageous for search." For example, this can be performed through a flow of pre-collection optimization, mid-consistency and quality assessment, and post-embedding and search optimization.

[0078] In one embodiment, the overall objective function can be defined as shown in the following mathematical formula 1.

[0079]

[0080] In mathematical formula 1 is the collector value (collection priority calculation formula), storable document quality, is embedding and search suitability, is the processing cost, and is a weight that reflects the importance of each element. This objective function aims not to “collect a large amount,” but to optimally secure only data that is ultimately high in value, reliable, and has good search performance, in a cost-effective manner.

[0081] Collection priority calculation formula using the first sub-formula It is equal to mathematical formula 2.

[0082]

[0083] In mathematical formula 2 is recency, is novelty, is source importance (importance by institution), is expected utilization, is overlapping risk, is the processing cost.

[0084] In the step of setting the sources and document types to be collected (S301), importance by institution , default weights by document type, collection budget , processing cost estimation methods, etc. are established.

[0085] In the step (S302) of accessing at least one of multiple public legal data providers to request list information, actual list information is collected, and the date, institution, type, page size, etc., for each document are secured. , It serves as the basis for calculation.

[0086] In the step (S303) of selecting new or updated targets by comparing with existing stored data, compared with existing stored data , Since it can be calculated, at this stage A value is calculated, and the collection target is selected first.

[0087] It makes it higher the closer the date of promulgation, decision, or response is, and It increases as the maximum similarity with existing documents decreases, and It may be set differently depending on the type of document, such as precedents, rulings, and authoritative interpretations, or the importance of each institution. It can reflect past search frequency, document type, query-answer logs, etc., and It can be defined as a duplicate possibility based on title, number, and body similarity. The actual adoption set is determined by optimization as shown in the following mathematical formula 3.

[0088]

[0089] In mathematical formula 3 is a processing budget that includes time, API call volume, CPU, OCR costs, etc. Equation 3 transforms the present invention from a simple sequential collection system into a budget-constrained value-maximizing collection system. Since collecting all public legal data is not always optimal and there are limits on call volume, OCR costs, parsing costs, and storage costs, target documents can be selected to maximize the sum of values ​​within the processing budget.

[0090] The second sub-formula is a combined formula for duplicate / update determination and URL consistency (Mathematical Formula 4).

[0091]

[0092] In the step (S303) of selecting new or updated targets by comparing with existing stored data, existing documents and candidate documents The overlap and update similarity between them is calculated using the above formula, and the existing document with the maximum value is considered the corresponding document. This value is the threshold. If it is greater than or equal to, it is determined to be an updated version; if it is less than or equal to, it is determined to be a new version. In the same manner, in the step (S309) of generating the source text access URL and search identifier, URL candidates The following consistency score can be calculated for (Equation 5).

[0093]

[0094] In mathematical formula 5 are the degree of agreement for title, number, institution, and date, respectively, and is a redirect or error response penalty.

[0095] The final URL can be determined as shown in the following mathematical formula 6.

[0096]

[0097] A URL is adopted as a valid URL only if it is greater than or equal to a threshold value. This structure can go beyond the existing "method of trying multiple candidate URLs" and evolve into a consistency optimization technique that numerically selects the optimal candidate among multiple candidate URLs.

[0098] The third sub-formula is a formula for determining the optimal threshold based on expected utility for whether to incorporate OCR. This aspect has significant practical impact in the field of legal data. While OCR is beneficial, its incorrect application can actually contaminate the original text. Therefore, a method based on expected utility, rather than a simple confidence cut-off, can be utilized.

[0099] In the step (S308) of supplementing the text by performing OCR on the image-containing area, the confidence score of the OCR result can be defined as shown in the following mathematical formula 7.

[0100]

[0101] In mathematical formula 7 is the character-unit average confidence, is word-level confidence, Legal terminology dictionary consistency, is the image noise score, It is a degree of layout collapse. Afterwards The probability After correction, the expected utility of OCR reflection Set it as. Here is the benefit obtained from reflecting correct OCR, This is the loss due to the reflection of misrecognition. The optimal condition for OCR reflection is Therefore, the OCR result can be inserted into the text only when the condition of the following mathematical formula 8 is met.

[0102]

[0103] Mathematical formula 8 is not set to an arbitrary constant value, but an optimal threshold value that is automatically determined based on the ratio of accurate reflection gain and misrecognition loss may be applied.

[0104] The fourth sub-formula is a document quality score formula that determines whether to save the final document (Mathematical Formula 9). The document quality score formula can be linked to the step (S310) of validating required fields and determining whether to save the document.

[0105]

[0106] In mathematical formula 9 is metadata integrity such as title, number, organization, date, etc. is URL consistency, is OCR reliability, is parsing success rate, is integrated text fidelity, is an error signal such as a missing field or abnormal format. The storage condition can be set as in Equation 10.

[0107]

[0108] In mathematical formula 10, the optimal threshold is the storage success profit and loss of incorrect storage Reflecting It can be induced in a manner similar to this. That is, if the cost of incorrectly stored legal documents contaminating the entire search and response system is high, the threshold is raised, and conversely, if the benefit of securing more documents is high even if some noise is allowed, the threshold is lowered.

[0109] The fifth sub-formula is a chunking optimization formula for the embedding and search index stage (Equation 11). The chunking optimization formula for the embedding and search index stage can be linked to the step of processing stored legal data into an embedding and search index (S312).

[0110]

[0111] In mathematical formula 11 Silver chunk length, The overlap length, is the semantic cohesion within a chunk, is context preservation, is search suitability, is indexing and search latency costs, is the cost of redundant storage. Equation 11 is directly connected to AI search. Rather than simple fixed-length splitting, a method that numerically determines semantic unit chunking suitable for the structure of legal documents can be used. The optimal chunking parameter can be determined by Equation 12.

[0112]

[0113] When mathematical formula 12 is included, the post-processing embedding and indexing pipeline becomes not a simple “creation of embeddings after chunk splitting,” but a semantic search preparation technique that optimizes chunk length and overlap length by simultaneously considering search performance and processing cost.

[0114] In one embodiment, the post-processing embedding and indexing pipeline unit (212) calculates a document-specific search suitability score for a plurality of chunk length and overlap length combinations, and can select the chunk length and overlap length that maximize the average or sum of the search suitability scores for the entire set of stored documents as the optimal chunking parameter.

[0115] Now, combining the above formulas, in the shear Select the target to collect, and in the interruption Only documents that can be saved are passed, and at the end... An embedding index is generated in a way that maximizes . That is, the entire system is summarized as in the following mathematical equation 13.

[0116]

[0117] In mathematical formula 13 Is It is the entire set of parameters including... In other words, a set of parameters that simultaneously increases new document recall rate, URL accuracy, OCR accuracy, and search sorting performance while lowering costs. It is a structure for finding. For example, this optimization can be performed using grid search with a set of legal documents for verification, coordinate descent, Bayesian optimization, etc.

[0118] FIG. 5 is a diagram illustrating the hardware configuration of a server according to one embodiment of the present invention.

[0119] Referring to FIG. 5, the server (10) may include a processor (110), memory (120), a transceiver (130), an input interface device (140), an output interface device (150), a storage device (160), an AI model (11), and a bus (170) that interconnects them. The processor (110) is the control entity of the present invention and can execute commands related to setting targets for collecting public legal data, requesting list information, selecting targets for new or updated data, collecting detailed document information, downloading original files or text conversion, generating structured data, generating integrated text, performing OCR, generating original URLs and search identifiers, validating data, storing data, embedding, and generating search indexes. The memory (120) may include ROM and RAM and can store the commands, temporary processing data, structured document data, OCR results, intermediate values ​​for generating embeddings, etc. The transceiver (130) can perform data transmission and reception with a plurality of public legal data providers, external storage, search index systems, or user terminals. The input interface device (140) can receive the administrator's setting input, collection condition input, or control command input, and the output interface device (150) can output processing results, logs, search results, or management screens. The storage device (160) can store programs, original files, structured data, logs, and model-related data non-volatilely. The AI ​​model (11) can be called or executed in conjunction by the processor (110) to perform intelligent processing such as OCR or embedding generation. The bus (170) supports the server (10) in performing the automatic collection and refinement method of public legal data of the present invention in an integrated manner by providing a transmission path for data and control signals between each of the above components.

[0120] Meanwhile, the method for automatic collection and refinement of public legal data according to one embodiment of the present invention may be implemented by one or more program instructions, and said program instructions may be loaded into the memory (120) of the server (10) and executed by the processor (110). When said program instructions are executed, the processor (110) may set the source and document type to be collected, access a plurality of public legal data providers to request list information, select new or updated targets by comparing with existing stored data, and collect detailed document information for the selected targets. In addition, the processor (110) may perform original file download or text conversion of non-standard documents, generation of structured document data, generation of integrated text, OCR on image-containing areas, generation of original access URLs and search identifiers, validation of essential fields, storage of structured data and original files, and processing of embedding and search index of the stored legal data. Accordingly, each step of the present invention is not limited to the fixed functions of the hardware itself, but can be realized by a program executed by the processor (110).

[0121] A method for automatically collecting and refining public legal data according to one embodiment may be implemented by a program stored on a computer-readable, non-transient recording medium. The recording medium may store one or more instructions executable by a processor, and when executed by a processor of a server, the instructions may be configured to perform the following: setting target sources and document types for collection; accessing public legal data providers and requesting list information; selecting new or updated targets through comparison with existing stored data; collecting detailed document information; downloading original files or text conversion of non-standard documents; generating structured document data; generating integrated text; performing OCR on image-containing areas; generating original URLs and search identifiers; validating essential fields and determining whether to save; saving structured data and original files; and processing the embedding and search index of the stored legal data.

[0122] The computer-readable recording medium may include ROM, RAM, flash memory, magnetic disk, optical disk, SSD, or a storage medium functionally equivalent thereto. Additionally, a program stored on the recording medium may be stored in the storage device (160) of the server (10), loaded into memory (120), and executed by a processor (110), and as a result of the execution, the server (10) may perform the method of automatic collection and refinement of public legal data described in FIG. 4.

[0123] The following supplements the definitions of the parameters used in each of the mathematical formulas described above.

[0124] is the total objective function representing the utility of the entire process of public legal data collection, refinement, and indexing. refers to individual public legal documents subject to processing. is a set of documents among the candidate documents that were actually selected for collection, storage, and indexing. is a set of parameters including weights, thresholds, chunking conditions, etc. used in the entire system. is document It is a score representing the collection priority or collection value. is document It is a quality score indicating whether it is suitable for final storage and utilization. is chunk length and overlap length Documents under conditions It is a score representing the search suitability of. is document It is the processing cost required to collect, refine, OCR, and index. These are weights that adjust the importance of collection value, document quality, search relevance, and processing cost, respectively. is document It is a score that indicates the recency or timeliness of. is document It is a score indicating the degree of high novelty compared to existing stored data. is a score indicating the importance of the data source or organization to which the document belongs. is document This is an expected utility score indicating the potential for use in follow-up searches, Q&A, analysis, etc. is document It is a duplication risk score indicating the likelihood of duplication with existing documents. is a weight that adjusts the reflection ratio of each detailed element in the collector value calculation formula. is the optimal set of documents selected to maximize the total value of collection under processing budget constraints. is the total processing budget or resource limit allowed during the entire collection and purification process. is a candidate document and existing saved documents It is a similarity score indicating the possibility of overlap or update between. refers to the comparison target document already stored in the database. is a weight that reflects the importance of similarity of title, number, body, and date. represents the title similarity between the candidate document and the existing document. represents the similarity between the document number or case number of a candidate document and an existing document. represents the similarity of the body content between the candidate document and the existing document. represents the similarity of date information between the candidate document and the existing document. is a duplicate / update similarity threshold for determining whether a candidate document is a new or updated document. is document and candidate URL It is a score indicating the consistency between them. refers to individual candidate URLs generated for accessing the original text. is the optimal URL with the maximum consistency score among multiple candidate URLs. is a weight that adjusts the reflection ratio of each element in the URL consistency score formula. is the page title and document pointed to by the candidate URL It refers to the degree of agreement between titles. is the page number information and document pointed to by the candidate URL It refers to the degree of agreement between the number information. is the organization information and document of the page pointed to by the candidate URL It refers to the degree of consistency between institutional information. is the date information of the page pointed to by the candidate URL and the document It refers to the degree of agreement between date information. is a penalty value for the extent to which a candidate URL causes a redirect, error response, or abnormal path. is document It is a score indicating the reliability of the OCR result. is a weight that adjusts the reflection ratio of each element in the OCR reliability calculation formula. represents the average confidence level per character in the OCR result. represents the word-unit average confidence in the OCR result. means the degree to which the OCR result matches a legal dictionary or legal context. is a score representing the noise level of the OCR target image. is a score indicating the degree of layout distortion or structural recognition error within the image. represents the expected utility when reflecting the OCR results in the actual text. is the probability or confidence probability that the OCR result is actually accurate. This is the benefit obtained by accurately reflecting the OCR results. is the loss that occurs by reflecting the OCR misrecognition results. is the optimal confidence threshold for determining whether to reflect the OCR results. is a score indicating the completeness and consistency of metadata such as title, number, organization, and date. is a URL quality score that indicates the extent to which the source URL points to the correct document. is a score indicating the quality or reflectability of the OCR result. is a score indicating the parsing success and stability of the source file or response data. is a score indicating the degree to which the integrated text is generated faithfully and searchably. is a penalty score representing the magnitude of error factors such as missing fields, format errors, and abnormal responses. is a weight that reflects the importance of each quality element in the document quality score formula. is the optimal quality threshold for determining whether to adopt the document as the final storage target. This is the benefit obtained by saving the document normally. This is the loss that occurs by storing low-quality documents. is the chunk length when splitting the document for embedding generation. is the overlap length between adjacent chunks. is a document under chunk length and overlap length conditions It is a score representing the embedding and search suitability of. is a weight that controls the importance of each element in the chunking optimization formula. is the document in the corresponding chunking condition It is a score representing the semantic cohesion within each chunk. is a context preservation score indicating the degree to which the preceding and succeeding context is maintained under the corresponding chunking conditions. is a search relevance score that indicates the extent to which the corresponding chunking condition contributes to search or query-answer performance. refers to the indexing or search delay cost resulting from the corresponding chunking condition. This refers to the cost of redundant storage or redundant calculation caused by overlap. is the optimal chunk length and optimal overlap length selected to maximize search suitability for the entire document set. It is the best overall set of parameters when comprehensively considering new document recovery rate, URL accuracy, OCR accuracy, search performance, and cost. These are weights that adjust the importance of new document recovery rate, URL accuracy, OCR accuracy, search performance, and cost terms, respectively. is the ratio of actual new documents that the system correctly collected. is the ratio of URLs among the generated URLs that actually point to the correct original text. It is an accuracy indicator that combines the precision and recall of OCR results. is a normalized discount cumulative gain indicator representing the ranking quality of search results. is the total cost consumed in the entire collection, refinement, OCR, embedding, and indexing process.

[0125] Although embodiments according to the technical concept of the present invention have been described above with reference to the attached drawings, those skilled in the art will understand that the present invention may be implemented in other specific forms without changing its technical concept or essential features. The embodiments described above should be understood as illustrative in all respects and not restrictive. Explanation of the symbols

[0127] 10: Server 11: AI Model 20: Database 30: User terminal 201: Collection Target Management Department 202: Heterogeneous Data Source Linkage Collection Unit 203: List Inquiry and Collection Target Selection Unit 204: Detailed Document Acquisition Section 205: Document Parsing and Refining Section 206: Integrated Text Generation Section 207: Image Character Recognition Processing Unit 208: Original URL Consistency Verification Section 209: Validity Check and Blacklist Processing Unit 210: Storage and File Upload Section 211: Search Identifier Generation Section 212: Post-processing Embedding and Indexing Pipeline 410: Input value 411: Image file for character recognition 412: Chunk-unit legal text containing semantic context 420: Output value 421: Recognized text string 422: High-dimensional embedding vector

Claims

Claim 1 A method for automatic collection and refinement of public legal data performed on a server comprises: a step of setting a source and document type to be collected; a step of accessing at least one of a plurality of public legal data providers and requesting list information based on the set source and document type to be collected; a step of selecting new or updated targets by comparing the list information with existing stored data; a step of collecting detailed document information for the selected targets; a step of downloading a source file corresponding to the detailed document information or converting a non-standard document into text; a step of generating structured document data from the detailed document information; a step of generating integrated text based on the structured document data; a step of supplementing the text by performing optical character recognition (OCR) on image-containing areas of the integrated text; a step of generating a source access URL and a search identifier based on the detailed document information, the structured document data, and the integrated text; a step of determining whether to save by verifying the validity of essential fields for the structured document data, the integrated text, the source access URL, and the search identifier; and a step of saving the structured document data and the source file, respectively, according to the result of the determination of whether to save. and includes the step of processing the stored legal data into an embedding and search index, wherein the step of selecting the new or updated target comprises each candidate document included in the list information Collector's value score regarding cast Calculating as, and pre-set processing budget The above collector value score within the range satisfying It includes selecting the new or updated target among the above candidate documents so that the sum of is maximized, and the above candidate documents is an individual public legal document identified from the above list information, and the above The above candidate document It is a recency score of 0 or more and 1 or less, calculated based on the time difference between the creation date, promulgation date, decision date, or update date of and the reference point, and the above is a novelty score of 0 or more and 1 or less calculated based on the result of comparison with the above-mentioned existing stored data, and the above The above candidate document It is a source importance score of 0 or more and 1 or less that indicates the importance of the collection target source to which belongs, and the above is an estimated utilization score of 0 or more and 1 or less calculated based on at least one of document type, organization type, historical search frequency, or reference frequency, and the above is a duplication risk score of 0 or more and 1 or less calculated based on at least one of the similarity of title, number, body text, or date, and the above is a processing cost of 0 or more calculated based on at least one of API call volume, download volume, OCR execution volume, and parsing operation volume, and the above inside is a weight that adjusts the reflection ratio of each item, and the above processing budget A method for automatic collection and refinement of public legal data, which is a budget value set based on at least one of the maximum API call volume, maximum processing time, maximum storage capacity, and maximum computation volume. Claim 2 In claim 1, the step of setting the source and document type to be collected comprises storing collection setting information for each of the plurality of public legal data providers, including an institution identifier, an institution name, key information used for generating integrated text, summary alternative field information, a document name extraction key, a document number extraction key, and setting information required for generating an original document access URL; and the step of accessing at least one of the plurality of public legal data providers based on the collection setting information and requesting list information, the step of collecting the detailed document information, the step of generating the structured document data, and the step of generating the original document access URL and search identifier; and the step of selecting new or updated targets comprises extracting a composite identification value by combining a serial number, document ID, case number, or two or more identification information included in the list information, comparing the extracted identification value with a corresponding identification value included in the existing stored data, and determining a new document or updated document as a selected target based on the comparison result, and collecting the detailed document information only for the selected targets. Claim 3 In claim 2, the step of generating the integrated text comprises: generating a header area including at least some of document name, institution name, document number, and date information from the structured document data; generating a body area including at least some of body text, summary, order, reason, background, appendix, and related laws from the structured document data; generating a footnote area including footnotes or additional explanations from the structured document data; and generating the integrated text by combining the header area, the body area, and the footnote area, wherein if a predetermined core field within the structured document data is empty, the integrated text is supplemented using the summary alternative field information; the step of generating the original text access URL and search identifier comprises generating a plurality of candidate original text access URLs and then determining the final original text access URL by comparing the title information or number information of the page indicated by the candidate original text access URL with the document information included in the detailed document information; the step of determining whether to save comprises determining whether to save by examining the existence and consistency of at least some of the document number, structured body data, integrated text, search identifier, and original text access URL; and the structured document data and the original text file, respectively A method for automatic collection and refinement of public legal data, wherein the saving step includes storing the structured document data in a database and uploading the original file to an external storage so as to link the storage path of the original file to the database. Claim 4 delete Claim 5 In claim 1, the step of supplementing the above text is each document corresponding to the image-containing area. OCR reliability score regarding cast Calculated as, the above OCR reliability score The probability value that the OCR result is accurate based on Calculating, and the benefits obtained by reflecting the above OCR results and the loss resulting from misapplying the above OCR results Optimal threshold based on cast Calculating as, and the above The above It includes reflecting the above OCR result in the above integrated text or above body only when there is an abnormality, and the above is a value between 0 and 1 representing the character-unit average confidence of the OCR result, and the above is a value between 0 and 1, representing the word-unit average confidence in the OCR result, and the above is a value between 0 and 1 indicating the degree of agreement between the OCR result and a pre-established legal terminology dictionary or legal context, and the above is a value between 0 and 1 indicating the degree of image noise, and the above is a value between 0 and 1 indicating the degree of distortion or layout recognition error of character placement within an image, and the above inside is a weight that adjusts the reflection ratio of each term, and the above is the above OCR reliability score A method for automatic collection and refinement of public legal data, which is a probability value between 0 and 1 inclusive, calculated by a monotonically increasing function with input. Claim 6 In claim 1, the step of generating the original text access URL and search identifier comprises each document corresponding to the detailed document information. and multiple candidate URLs URL consistency score for each cast Calculated as, and the above URL consistency score The step of confirming the candidate URL with the maximum value as the final source access URL, and the said candidate URL is a candidate link for accessing the source text generated by applying multiple offset values ​​or multiple URL generation rules, and the above is the above candidate URL Title information of the page indicated by [the document] It is a value between 0 and 1 inclusive indicating the degree of agreement between the title information, and the above is the above candidate URL Page number information indicated by A and the above document It is a value between 0 and 1 inclusive representing the degree of agreement between the number information, and the above is the above candidate URL The institution information on the page indicated by A and the above document It is a value between 0 and 1 indicating the degree of agreement between the institution information, and the above is the above candidate URL Date information of the page indicated by [the document] It is a value between 0 and 1 inclusive indicating the degree of agreement between the date information, and the above is the above candidate URL is a penalty value between 0 and 1 indicating the degree to which it causes a redirect, error response, or abnormal path, and the above inside A method for automatic collection and refinement of public legal data, which is a weight that adjusts the reflection ratio of each clause. Claim 7 In claim 1, the step of processing into the embedding and search index comprises the stored legal data chunk length and overlap length Each document divided according to Search relevance score for cast Calculated as, and the above search relevance score for the entire set of stored documents Optimal chunk length to maximize the average or sum and optimal overlap length Select and the optimal chunk length above and the optimal overlap length above It includes chunking the stored legal data based on [the data], generating an embedding vector, and loading it into a search index, and [the chunk length] is a value representing the number of characters or tokens included in each chunk, and the overlap length is a value representing the number of characters or tokens maintained to be duplicated between two adjacent chunks, and the above is the above chunk length and the above overlap length Documents under conditions It is a value between 0 and 1 indicating the semantic cohesion within each chunk, and the above is the above chunk length and the above overlap length It is a value between 0 and 1 inclusive indicating the degree to which the preceding and succeeding context is maintained under the condition, and the above is the above chunk length and the above overlap length A value between 0 and 1 inclusive representing the degree to which it contributes to search or query-response performance in the condition, and the above is the above chunk length and the above overlap length A value greater than or equal to 0 representing the indexing or search delay cost due to the condition, and the above is the above overlap length A value greater than or equal to 0 representing the cost of duplicate storage or duplicate calculation caused by, and the above inside A method for automatic collection and refinement of public legal data, which is a weight that adjusts the reflection ratio of each clause. Claim 8 In a public legal data automatic collection and refinement system, a server including an AI model and a database; The server includes a user terminal connected to communicate with the server, wherein the server is configured to automatically collect public legal data from a plurality of public legal data providers, generate structured document data and integrated text for the collected public legal data, perform intelligent processing of the public legal data using the AI ​​model, and store the structured document data and the integrated text in the database; wherein the user terminal is configured to transmit a request for inquiry, search, or management of public legal data to the server, and output refinement results, search results, document body, or original text access information provided by the server; wherein the server is configured to connect to at least one of the plurality of public legal data providers to request list information, select new or updated documents among the candidate documents included in the list information, and collect detailed document information regarding the new or updated documents; wherein for each candidate document included in the list information, the server is configured to provide a recency score based on the time difference between the date of creation, date of promulgation, date of decision, or date of update of the candidate document and a reference point, a novelty score based on the result of comparison with existing stored data, a source importance score indicating the importance of the collection target source to which the candidate document belongs, document type, and institution A collection value score is calculated using an estimated utilization score based on at least one of type, past search frequency, or reference frequency; a duplicate risk score based on at least one of similarity of title, number, body, or date; and a processing cost based on at least one of API call volume, download volume, OCR execution volume, and parsing operation volume; and a new or updated target is selected from the candidate documents such that the sum of the collection value scores is maximized within a range satisfying a preset processing budget, and the server is configured toFor each document corresponding to the detailed document information and each of the multiple candidate URLs, a URL consistency score is calculated using the degree of agreement between the title information of the page indicated by the candidate URL and the title information of the document, the degree of agreement between the number information of the page indicated by the candidate URL and the number information of the document, the degree of agreement between the institution information of the page indicated by the candidate URL and the institution information of the document, the degree of agreement between the date information of the page indicated by the candidate URL and the date information of the document, and a penalty value indicating the extent to which the candidate URL causes a redirect, error response, or abnormal path; and the candidate URL with the maximum URL consistency score is determined as the final original document access URL. The candidate URL is an original document access candidate link generated by applying multiple offset values ​​or multiple URL generation rules. For each document in which the stored public legal data is divided according to chunk length and overlap length, the server considers the semantic cohesion within each chunk under the conditions of the chunk length and overlap length, the degree to which the preceding and succeeding context is maintained, the degree to which it contributes to search or query response performance, indexing or search latency costs, and duplicate storage or duplicate calculation caused by the overlap length. A public law data automatic collection and refinement system configured to calculate a search suitability score using cost, select an optimal chunk length and an optimal overlap length such that the average or sum of the search suitability scores is maximized for the entire set of stored documents, and, based on the optimal chunk length and the optimal overlap length, chunk the stored public law data and generate an embedding vector to load into a search index. Claim 9 A public legal data automatic collection and refinement system according to claim 8, wherein the AI ​​model is configured to receive an image file for character recognition and output a recognized text string, or to receive a chunk-unit legal text containing semantic context and output a high-dimensional embedding vector, and the server is configured to reflect the recognized text string in the integrated text or body and to generate a search index for the public legal data using the high-dimensional embedding vector. Claim 10 A public legal data automatic collection and refinement system according to claim 9, wherein the AI ​​model comprises as training data pairs of image files for character recognition extracted from legal documents and correct text strings corresponding to the image files for character recognition, and is trained to update the parameters of the AI ​​model so as to reduce the loss value between the output of the AI ​​model, which receives the image files for character recognition and outputs a predicted text string, and the correct text strings. Claim 11 A public legal data automatic collection and refinement system according to claim 9, wherein the AI ​​model is trained to update the parameters of the AI ​​model according to a loss function that decreases the distance between embedding vectors for the positive learning pairs and increases the distance between embedding vectors for the negative learning pairs, and comprises chunk pairs indicating the same document, the same event, or the same statute among chunk-unit legal texts containing semantic context as positive learning pairs and chunk pairs indicating different documents, different events, or different statutes as negative learning pairs.