Ai inference-based method for collecting judgments and a system utilizing the same
Patent Information
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- LAW COMPANY
- Filing Date
- 2026-03-19
- Publication Date
- 2026-08-03
Smart Images

Figure 112026033770585-PAT00100_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to a method for collecting judgment documents based on AI inference and a system utilizing the same. Background Technology
[0002] Various judicial documents, such as judgments, decisions, and orders, serve as core resources for legal research, dispute response, precedent analysis, the establishment of internal knowledge bases, and legal information services for the public. However, traditionally, the channels for securing these documents were not unified but fragmented across multiple platforms, including the Supreme Court's online viewing service, copy request procedures, comprehensive legal information services, and Constitutional Court case searches. Consequently, it was difficult to collect and manage necessary documents in a consistent manner. In particular, some documents required paid viewing or separate application procedures, while others were provided only in preview or HTML format. Furthermore, since provision methods, search conditions, and metadata structures varied depending on the type of case or court, users had to repeatedly perform the processes of searching, verifying, and saving individually, tailored to the screen layout and input format of each service.
[0003] Furthermore, conventional judgment collection tasks did not end with merely downloading documents; they involved additional work such as verifying duplicates, organizing metadata like case numbers, court names, and judgment dates, post-processing based on original text formats, and converting the data into a service-ready database structure. However, in existing methods, these follow-up processes were often limited to manual work or individual scripts, leading to increased workloads during large-scale collection and making it prone to processing omissions or format inconsistencies. Moreover, original judgment texts frequently contain a mixture of images, tables, footnotes, attachments, and unstructured text, requiring significant time and manpower to refine them into structured data suitable for search and analysis. Consequently, there is a continuously raised technical need for the ability to reliably collect judgment-related documents from various sources and systematically organize and manage them in a format suitable for subsequent utilization. The problem to be solved
[0004] The present invention resolves the inefficiency and instability that occur during the process of securing judicial documents, such as judgments, decisions, and orders.
[0005] Traditionally, since multiple acquisition channels—such as internet viewing, copy requests, public case law searches, and preview information—had different authentication methods, search conditions, provision formats, and subsequent processing procedures, users had to repeatedly search and verify the same case separately for each channel. This process frequently resulted in omissions, duplications, delays, and management errors. Furthermore, because collected documents existed in various formats such as PDF, HTML, and summary text, additional organization and refinement were required for subsequent searching, analysis, database reflection, and service provision; however, there was a problem in consistently performing these tasks.
[0006] Accordingly, the present invention aims to promote the automation, standardization, and improvement of operational efficiency in the overall task of collecting judgments by systematically linking multiple judgment acquisition paths to stably secure judicial documents through a path suitable for the collection target, and by organizing and storing the secured original text or preview information in a data format that can be subsequently utilized.
[0007] The problems of the present invention are not limited to those mentioned above, and other unmentioned technical problems will be clearly understood by a person skilled in the art from the description below. means of solving the problem
[0008] An AI inference-based judgment collection system according to an embodiment of the present invention for solving the above problem comprises: a data server that provides legal data such as judgments; a user terminal that receives a user's request to collect judgments and outputs a judgment collection result; and a collection server that is connected to communicate with the data server and the user terminal and includes an AI model that analyzes request information transmitted from the user terminal and previously secured data to derive a direction for the collection or processing of judgments, and a database that stores original judgment text, metadata, collection history, user request history, AI inference result, duplicate determination result, and refined structured data.
[0009] A method for collecting judgment documents based on AI inference performed on a server including an AI model according to an embodiment of the present invention for solving the above problem comprises: a multi-stage automatic authentication and session automatic recovery step in which an authentication token is extracted from the main page of a web portal provided by an external server, said authentication token is set in a cookie, and an access session for collecting judgment documents is formed by performing CAPTCHA verification and login processing, and if session expiration is detected during the execution of a subsequent step, said authentication token extraction, cookie setting, CAPTCHA verification, and login processing are re-executed to automatically recover said access session; an automatic court document collection eligibility determination step in which a case code is extracted from a case number based on said access session, a verification of whether the case is subject to exclusion and whether the portal's internal code can be converted, and a court name and case code for collectible cases are converted into an internal identification system to generate request parameters for subsequent collection or application; an API-RPA hybrid payment and automatic recovery step in which, as a result of performing a judgment document search or application using said request parameters, paid viewing or fee payment is required, payment is performed by combining API-based status inquiry and user interface automation, a recovery action is performed in case of payment failure, and the judgment document or judgment document-related data is acquired after payment is completed; said judgment document or If the data related to the judgment is preview text or partially omitted legal text, it is separated into multiple fragments based on ellipsis symbols, the completion status of each fragment is determined, extended keywords are derived from incomplete fragments to perform a re-search, and new fragments and existing fragments are combined, but if session expiration is detected during the re-search, the above multi-step automatic authentication and session automatic recovery steps are re-executed to continue the re-search, and if it remains incomplete even after combining, the above extended keyword derivation, re-search, and combination are repeated in an iterative restoration step.and includes a data migration step that determines identity by comparing the judgment obtained in the API-RPA hybrid payment and automatic recovery step or the judgment restored in the iterative recovery step with other judgments obtained from multiple collection channels, changes the service availability status or replaces it with new data according to source priority, and updates the reference ID of the original table to the new ID of the integrated table. Effects of the invention
[0010] According to an embodiment of the present invention, judicial documents having different acquisition procedures and provision formats can be collected and managed within a single system, thereby significantly reducing the manpower and time required to secure judgments. Furthermore, since the processes of authentication, search, payment, copy request, download, storage, and subsequent refinement are processed continuously, the burden of personnel having to access individual sites to perform repetitive tasks, as was the case in the past, can be reduced. Additionally, operational errors such as session expiration, payment failure, duplicate collection, and format inconsistencies can be responded to more stably. Moreover, by organizing collected documents along with metadata and structuring and refining them according to the original text format to reflect them in a database, it is possible to build judgment data assets suitable for searching, analysis, service provision, and internal knowledge storage, going beyond mere file storage. In particular, even partial preview information or unstructured HTML original texts can be organized into a form suitable for subsequent use, thereby expanding the scope of collectible judgment information and enhancing the completeness and usability of service data.
[0011] The effects according to the embodiments are not limited to those exemplified above, and a wider variety of effects are included in this specification. Brief explanation of the drawing
[0012] FIG. 1 is a schematic diagram illustrating the components of an AI inference-based judgment collection system according to one embodiment of the present invention. FIG. 2 is a diagram illustrating the functional elements of a collection server according to one embodiment of the present invention. FIG. 3 is a flowchart illustrating a method for collecting judgment documents according to one embodiment of the present invention. Figure 4 is a diagram illustrating the function of the AI model in the iterative restoration step of Figure 3. FIG. 5 is a diagram illustrating the hardware configuration of a collection server according to one embodiment of the present invention. Specific details for implementing the invention
[0013] The advantages and features of the present invention and the methods for achieving them will become clear by referring to the embodiments described below in detail together with the accompanying drawings. However, the present invention is not limited to the embodiments disclosed below but may be implemented in various different forms. These embodiments are provided merely to ensure that the disclosure of the present invention is complete and to fully inform those skilled in the art of the scope of the invention, and the present invention is defined only by the scope of the claims.
[0014] Although terms such as "first," "second," etc., are used to describe various components, it goes without saying that these components are not limited by these terms. These terms are used merely to distinguish one component from another. Therefore, it goes without saying that the first component mentioned below may be the second component within the technical scope of the present invention. A singular expression includes a plural expression unless the context clearly indicates otherwise.
[0015] As used below, the term "judgment" may be used for convenience to encompass judgments, decisions, orders, and judicial documents of a similar format. Accordingly, in this specification, the terms judgment, judgment, decision, order, or judicial document may be used interchangeably depending on the context.
[0016] 'AI inference' may refer to an information processing process that derives collection targets, collection priorities, processing directions, restoration directions, or search term candidates based on input request information, contextual information, or previously acquired data, and may include rule-based processing, statistical-based processing, machine learning-based processing, and combinations thereof.
[0017] 'Incomplete legal text' may refer to a legal text in which some preceding or succeeding sections are missing, parts of the text are replaced by ellipses, or the overall context is incomplete.
[0018] 'Incomplete fragment' may refer to a fragment among each text fragment obtained by separating an incomplete legal text based on ellipsis symbols, in which the initial or final completeness is not satisfied.
[0019] 'Extended keywords' are derived from the context of the preceding or succeeding part of an incomplete fragment and may refer to re-search terms for exploring a more complete text that includes the omitted part.
[0020] "Portal internal identification system" may refer to internal codes, registration names, identification values, or combinations thereof required by an external portal or application system for case search, application, or download processing, and court names, case codes, and case numbers may be converted to conform to the said portal internal identification system.
[0021] Hereinafter, embodiments of the present invention will be described in detail with reference to the attached drawings. Identical or similar reference numerals are used for identical components in the drawings.
[0022] FIG. 1 is a schematic diagram illustrating the components of an AI inference-based judgment collection system according to one embodiment of the present invention.
[0023] Referring to FIG. 1, an AI inference-based judgment collection system according to one embodiment of the present invention may be implemented as a server-client structure including a collection server (10), a data server (20), and a user terminal (30). More specifically, the user terminal (30) receives a collection request from a user and transmits it to the collection server (10), and the collection server (10) performs search, identification, inference, refinement, and storage of judgments to be collected using an AI model (11) and a database (12) included internally, and the data server (20) may be connected to the collection server (10) for data communication as an external or separate management server that provides legal data such as judgments. That is, the present system has a structure that integrally performs a series of operation flows, such as request input from the user terminal (30), inference and processing at the collection server (10), securing source data from the data server (20), and re-providing the processing results.
[0024] The collection server (10) is a central component of the present invention and is an electronic device that performs the entire method of collecting judgment documents. The collection server (10) may be physically implemented as one or more computing servers, virtual servers, cloud instances, or container-based processing nodes, and functionally performs functions such as managing collection requests, identifying collection targets, accessing external legal data sources, storing collection results, and generating user responses. For example, the collection server (10) may receive request information regarding case number, court name, judgment date, case name, party information, search keywords, or collection scope from a user terminal (30), and based on the received request information, determine which data source to access first, what format of data to request, and whether the same or similar case has already been secured. Additionally, the collection server (10) may post-process the original text or metadata of a judgment document received from a data server (20), convert it into searchable structured data, store it in a database (12), and provide the processing result to the user terminal (30).
[0025] The AI model (11) is an inference engine that is included in and operates within the collection server (10). It supports the collection server (10) in not merely operating on pre-determined rules, but in deriving a more suitable direction for collection or processing by analyzing the input request information and the state of previously secured data. For example, the AI model (11) receives input values such as search terms, case numbers, case types, court information, previous collection history, and document format information transmitted from the user terminal (30), and can infer data sources with a high probability of containing the judgment, priority collection targets, whether subsequent refinement is necessary, and the possibility of duplication or structuring.
[0026] As another embodiment, the AI model (11) may operate by identifying key areas such as case names, rulings, orders, reasons, and reference provisions within the collected original text, or by deriving supplementary candidates based on surrounding context and previously stored data when some information is missing.
[0027] Meanwhile, the AI model (11) can operate as an inference module that identifies key semantic words included in incomplete legal text fragments and generates expanded keywords. In one embodiment, the AI model (11) receives a text of a predetermined proportion of the front or rear portion of an incomplete fragment as input, performs morphological analysis, extracts nouns among the morphemes tagged with parts of speech, and then removes stop words and duplicate words to select proper nouns as expanded keywords. According to this configuration, keywords valid for re-searching in Korean legal sentences can be derived more accurately, thereby increasing the possibility of restoring omitted parts of the judgment.
[0028] That is, the AI model (11) can be understood as a key processing means to improve the judgment accuracy and automation level of the collection server (10).
[0029] The database (12) is a storage means linked with the collection server (10) to store the original text of the judgment, metadata, collection history, user request history, AI inference results, duplicate determination results, refined structured data, etc. The database (12) can be implemented as a relational database, a document database, a vector database, or a combination thereof. For example, the database (12) can store the case number, court name, date of judgment, judgment file path, hash value, full text of the original text, summary data, classification results of the AI model (11), the last collection time, etc. Through this, the collection server (10) can determine whether to process the same request repeatedly, prevent duplicate storage by comparing existing stored data with newly received data, and perform subsequent structuring or re-analysis on the already stored data.
[0030] The data server (20) may be an external data source or an internal linkage server that provides legal data such as judgments. For example, the data server (20) may be implemented as a group of servers that stores and manages original judgments, basic case information, summary of rulings, decisions, orders, public materials by court, or other legal documents. The collection server (10) may connect to the data server (20) to transmit a search request and receive a judgment file, original text, HTML document, PDF document, metadata, or download path information in response. As another embodiment, the data server (20) may be a set of multiple data sources rather than a single server, and the collection server (10) may access these multiple data sources sequentially or in parallel to secure data with a high probability of successful collection. Therefore, the data server (20) can be understood as a target layer for the collection server (10) to acquire external legal data, rather than a simple file storage.
[0031] The user terminal (30) is a component for the user to interact with the system. The user terminal (30) can be implemented as a desktop computer, laptop, tablet, smartphone, workstation, etc., and can communicate with the collection server (10) through a web browser, a dedicated client program, or a work management screen. At the user terminal (30), the user can input a collection request for a specific case, specify a collection scope, or view and download collection results. For example, the user can request the collection of judgment documents by entering a case number or keyword through the user terminal (30), and after processing the request, the collection server (10) can return to the user terminal (30) whether the collection was successful, a list of secured documents, structured judgment information, or whether a retry is required. That is, the user terminal (30) functions as an input interface and result verification interface of the system, going beyond a simple display means.
[0032] In one embodiment, request information input from a user terminal (30) is transmitted to a collection server (10). The collection server (10) refers to existing history and internal status stored in a database (12), analyzes the request information through an AI model (11), and derives an appropriate collection strategy. Subsequently, the collection server (10) communicates with a data server (20) to collect necessary judgments or related legal data, refines or structures the collected data, and stores it in the database (12). The stored results or summarized processing results are then provided back to the user terminal (30). Thus, by combining the data flow of "User terminal (30) → Collection server (10) → Data server (20)" and the result flow of "Data server (20) → Collection server (10) → Database (12) → User terminal (30)", the present invention can be implemented as a system in which the collection, inference processing, storage, and provision to the user of judgments are integrated. In particular, since the AI model (11) operates in conjunction with the database (12) inside the collection server (10), it has technical significance in that it is possible to go beyond simple collection to determine request suitability, determine collection priority, improve data consistency, and increase subsequent usability.
[0033] FIG. 2 is a diagram illustrating the functional elements of a collection server according to one embodiment of the present invention.
[0034] Referring to FIG. 2, the collection server (10) may be implemented as a complex functional server that not only performs the function of simply accessing an external site to download a judgment file, but also organically performs a series of processes including authentication, search, application, payment, restoration, and structuring in response to multiple acquisition paths. To this end, the collection server (10) may include a judgment collection module (210), an automatic authentication module (220), a hybrid payment automation module (230), an incomplete legal text restoration engine (240), a multi-layered verification and court code dual mapping system (250), and an automatic judgment structure decomposition engine (260). Each functional element may operate independently, but in actual implementation, it is desirable to form a single integrated collection pipeline by interoperating according to the type of request and the state of the collection target.
[0035] First, the judgment collection module (210) is a core module responsible for the external collection function of the collection server (10). Based on user requests or pre-set collection conditions, the judgment collection module (210) accesses multiple data sources, such as Supreme Court internet viewing, copy request, preview, comprehensive legal information, Constitutional Court case search, and court public bulletin board, and can obtain the original judgment text, preview text, metadata, attached files, or download path from each data source. For example, when a specific case number is entered, the judgment collection module (210) may first attempt a search through a paid internet viewing path, switch to a copy request path if it is difficult to obtain through that path or if it is not subject to internet viewing, and if a public case exists, it may be configured to obtain the HTML original text through comprehensive legal information or the Constitutional Court portal. Additionally, the judgment collection module (210) can perform detailed operations such as page-by-page traversal of search results, calling a detailed inquiry API, downloading the original text, collecting attached PDFs, retrieving image files, uploading to GCS or other external storage, and generating metadata. That is, the judgment collection module (210) can be described as an interface layer that directly connects with an external data source at the forefront of the collection server (10).
[0036] The automatic authentication module (220) is a functional element that forms and maintains a session and authentication state so that the judgment collection module (210) can reliably access an external portal. Since many judgment provision sites require procedures such as non-member or associate member-based login, cookie setting, granting of authentication tokens, verification of anti-automatic input characters, and mobile phone number authentication, it is difficult to reliably collect with only simple HTTP requests. Accordingly, the automatic authentication module (220) can extract an authentication token, such as certToken, from the portal main page, set it as a session identifier such as a scotkn cookie, and configure request headers to form a state similar to normal browser access. In addition, the automatic authentication module (220) can pass the anti-automatic input procedure by receiving CAPTCHA image data, deriving an authentication character through an internal character recognition function or an external recognition server, and transmitting it to a verification API. Furthermore, it can be configured to include basic login API calls, mobile phone number-based authentication, transmission of non-member login information, determination of session expiration, and automatic re-login upon expiration. Therefore, the automatic authentication module (220) corresponds to the preprocessing module of the judgment collection module (210), and has functional priority in that if this module does not operate normally, all processing such as subsequent search, loading into a shopping cart, requesting a copy, and downloading can be blocked.
[0037] The hybrid payment automation module (230) is an element for handling cases where paid internet viewing or fee payment is involved in the process of collecting judgments. For some judgments, even if searching itself is possible, loading into a shopping cart and card payment may be required for actual viewing or downloading, and for some copy requests, a separate account transfer may be required in accordance with the fee payment notice sent by the court. To process these heterogeneous payment procedures in a unified manner, the hybrid payment automation module (230) may have a combined structure of browser-based UI operation and data-based subsequent processing. For example, in the card payment stage, after checking for the existence of unpaid items by querying the shopping cart API, the selection of a payment button, input of card information, processing of the ISP security module, confirmation of payment completion, and browser cleanup can be automatically performed using an RPA tool. In the event of a payment failure, it may be configured to include the termination of the ISP process, reinstallation, retry, and failure notification. Meanwhile, in the subsequent step of the copy application, the HTML table of the payment request email is parsed to extract the bank name, account number, amount, and receipt number; an Excel file for account transfer requests and a PDF for transfer verification are generated and delivered to the person in charge; and the final PDF is received from the download URL included in the delivery completion email. That is, the hybrid payment automation module (230) is an intermediate operating layer that handles both the electronic payment type procedure and the administrative fee payment type procedure, and is responsible for the function of converting the target cases secured by the judgment collection module (210) into a state where they can be actually secured.
[0038] The multi-layered verification and court code dual mapping system (250) is a preprocessing and judgment layer designed to ensure the accuracy and processability of the collected data. External portals typically do not use the court name or case number formats recognized by users as they are, but require internal formats such as internal court codes, case code codes, registration names, civil / criminal distinction values, and digit-corrected serial numbers. Accordingly, the multi-layered verification and court code dual mapping system (250) first separates the year, case code, and serial number from the case number and can determine whether the case code is criminal, civil, or of other types. Additionally, it checks whether the case code is excluded from application according to precedents or system policies, verifies whether it exists in the list of registered case codes, and if it is an unregistered code, it can treat it as a failure or branch it to a manual review target. Along with this, it can be configured to perform a first mapping between a user-friendly court name and the portal's internal court code, and a second mapping between the court registration name and registration code required by the copy application system. In other words, dual mapping refers to a structure that does not merely replace a single court identification value with a single code, but rather converts it into a separate set of codes corresponding to two or more types of court identification systems required by different external services. This functional element performs the role of aligning search parameters between the judgment collection module (210) and the automatic authentication module (220), and acts as a gateway to screen the legality and feasibility of a case before proceeding to the hybrid payment automation module (230) or the copy application processing flow.
[0039] The incomplete legal text restoration engine (240) is a functional element that restores or supplements text in situations where only preview information or partially omitted legal text is provided, rather than the entire judgment. Since the Supreme Court portal, etc., may disclose only parts of the order and reasons and provide a preview in which the middle or preceding and succeeding sections are omitted as "...", it is difficult to secure serviceable case text through simple collection alone. Accordingly, the incomplete legal text restoration engine (240) first separates the text into fragments based on ellipsis symbols and can determine the start completeness (start_flag) and end completeness (end_flag) for each fragment. For example, if the beginning of a fragment starts with an ellipsis symbol, it is determined that the front part is omitted, and if the end does not conclude in the form of a judgment ending or an order ending, it is determined that the latter part is omitted. Then, morphological analysis is performed on parts of the preceding or succeeding sections of the incomplete fragments to extract nouns, and after removing stop words and duplicate words, core keywords suitable for extended search can be derived. The derived keywords can be transmitted back to the judgment collection module (210) as subsequent search conditions and used for collecting additional fragments. Furthermore, when multiple fragments are secured, the possibility of connection between fragments is determined based on overlapping strings of a certain length or longer, and they can be combined into a single continuous text using patterns such as identical, inclusion, and front-to-back joining. That is, the incomplete legal text restoration engine (240) is important in that it functions not as a post-processor that performs simple organization after collection, but as a feedback engine that induces the re-search operation of the judgment collection module (210).
[0040] The judgment structure automatic decomposition engine (260) is a functional element that decomposes the acquired original text into structural data suitable for subsequent search, analysis, and service. The original text of the judgment exists in various formats such as PDF, HTML, or plain text, and may contain a mixture of case expressions, case names, ruling matters, summary of judgment, reference provisions, reference precedents, orders, claims, reasons, judges, attachments, footnotes, tables, images, etc. Accordingly, the judgment structure automatic decomposition engine (260) may first apply different preprocessing strategies depending on the format of the original text. For example, for original texts in the comprehensive legal information series, images and tables may be replaced with placeholders and pure text may be extracted, and for original texts in the Constitutional Court series, signature image removal, footnote extraction, table organization, "HTML → text conversion," and spacing normalization may be performed. Afterward, using a regular expression-based pattern or a rule-based parser, the court name, judgment date, case number, and judicial division can be extracted from the case expression in line 1, the case name and publication information in line 2 can be separated, and the body can be sequentially structurally decomposed into areas such as ruling matters, summary of decision, parties, order, reasons, and judges. Additionally, footnotes and tables within the judgment can be restored while maintaining the correspondence between the placeholder and the original, and stored in the form of separate fields or reference tags within the body. Thus, the judgment structure automatic decomposition engine (260) is a core analysis layer that ultimately generates a serviceable cases table or a corresponding storage structure, and the derived structural information can be reused for duplicate determination, case search, summary extraction, legal research support, etc.
[0041] Each of these functional elements has a sequential and cyclic connection. For example, the automatic authentication module (220) establishes access rights to an external portal to enable the initial collection of the judgment collection module (210), and the multi-layered verification and court code dual mapping system (250) improves the search accuracy of the judgment collection module (210) by converting the case number and court name of the collection target into an internal format required by the external system. Subsequently, if paid acquisition is required, the hybrid payment automation module (230) intervenes to perform subsequent procedures such as checking the shopping cart status, card payment, or fee payment, and as a result, supports the judgment collection module (210) in securing the final PDF or related file. Meanwhile, if only a preview is secured, the partial text collected by the judgment collection module (210) is transmitted to the incomplete legal text restoration engine (240), and the expanded keywords generated by the restoration engine (240) are transmitted back to the judgment collection module (210) to induce additional searching. Subsequently, the original text or restored text is input into the automatic judgment structure decomposition engine (260) to extract case information and body structure, and the extraction result can be stored in a database (12) inside the collection server (10) or in an external storage. That is, each functional element of FIG. 2 is not a simple parallel arrangement, but forms a closed-loop processing structure in which re-authentication, re-search, and re-combination are repeated as needed while forming a multi-stage pipeline of “authentication → verification / mapping → collection → payment or restoration → structuring.”
[0042] In terms of implementation examples, each functional element may be implemented as a software module, a microservice, a scheduler-based batch job, a containerized unit of work, or a combination thereof. For example, the judgment collection module (210) may be implemented as a Scrapy-based crawler or API call service, the automatic authentication module (220) may be implemented as a session management class and a CAPTCHA recognition call unit, and the hybrid payment automation module (230) may be implemented as a browser automation script, an email parser, and a subsequent file generator. The incomplete legal text restoration engine (240) may be composed of text separation logic, completeness determination logic, a morphological analyzer, and a fragment combiner, the multi-layered verification and court code dual mapping system (250) may be composed of a case code dictionary, a court code mapping table, a set of rule exclusion rules, and a format normalization module, and the judgment structure automatic decomposition engine (260) may be implemented as an HTML preprocessor, a regular expression parser, a postprocessor, and a migration processor. However, this implementation method is merely an example, and the present invention can be modified into various forms, such as hardware circuits, firmware, scripts, artificial intelligence assistance modules, or distributed processing environments, as long as they perform the same function.
[0043] Ultimately, it is reasonable to understand that the collection server (10) is not a single-function device for acquiring judgments, but rather a functional set that integrates access control to external legal data sources, case-unit verification, payment procedure processing, restoration of incomplete information, and structuring of original texts. In particular, the important technical significance of this embodiment is that judgment collection module (210), automatic authentication module (220), hybrid payment automation module (230), incomplete legal text restoration engine (240), multi-layer verification and court code dual mapping system (250), and judgment structure automatic decomposition engine (260) are interconnected, thereby enabling judgments having different delivery methods and formats to be secured, refined, and stored in a stable and consistent form.
[0044] FIG. 3 is a flowchart illustrating a method for collecting judgment documents according to one embodiment of the present invention.
[0045] Referring to FIG. 3, the method for collecting judgment documents according to an embodiment of the present invention is not merely a procedure of accessing a specific site and downloading a file, but consists of a series of interconnected processes that automatically determine the collection target in a court portal environment requiring authentication, perform payment if necessary, repeatedly restore incompletely disclosed legal texts, and migrate results obtained from multiple sources into integrated data. That is, the present method may include a multi-stage automatic authentication and session automatic recovery step (S310), an automatic determination of eligibility for court document collection step (S320), an API-RPA hybrid payment and automatic recovery step (S330), an iterative restoration step (S340), and a data migration step (S350). Each step may be performed sequentially, but in actual implementation, it may operate in a closed-loop structure where the result of the preceding step determines the processing condition of the subsequent step, and if authentication expiration or restoration failure is detected in the subsequent step, the preceding step is called again. Therefore, it is preferable to understand each step of FIG. 3 not as an independent unit function, but as a processing step organically combined to stabilize the entire judgment document collection pipeline.
[0046] First, the multi-step automatic authentication and session automatic recovery step (S310) is a step for the collection server (10) to establish and continuously maintain a state in which it can access an external court portal or a legal information provision server. Unlike general public APIs, the court portal may require multiple authentication procedures, such as main page access, cookie setting, CAPTCHA authentication, user information-based login, and mobile phone authentication; therefore, this step performs the function of mechanically reproducing these procedures. Specifically, the collection server (10) may first request the HTML of the main page of the web portal to extract an authentication token. The authentication token may be, for example, a certToken, which can be used as a value that guarantees the validity of subsequent authentication requests. Next, the collection server (10) may set the authentication token in a browser cookie or session storage; for example, by storing it in a cookie named scotkn, the portal may recognize the collection server (10) as a legitimate session participant. Subsequently, the collection server (10) receives a CAPTCHA image and transmits the received image data to an external character recognition server or an internal recognition engine to obtain a reading result. At this time, a verification API is called using the verification string provided along with the CAPTCHA image, and if the reading result matches, the process proceeds to the next login step; if the reading result does not match, a new CAPTCHA image is received and a retry can be attempted. The retry can be performed using an exponential backoff method rather than simple repetition, thereby preventing portal overload in the event of consecutive failures and increasing the probability of successful authentication. Afterward, the collection server (10) transmits user identification information and an authentication token to the login API, and if necessary, additionally transmits an encoded mobile phone number and an authentication token to the phone number authentication API to form a login state that enables shopping cart viewing, payment, downloading, etc. In particular, an important feature of this step is the session recovery function.That is, the collection server (10) monitors the value of a specific session state field, such as userCertfLgnYn, in the result of a subsequent API call, and if the value indicates a non-login state, determines that the session has expired and can automatically re-execute authentication token reissuance, cookie reset, CAPTCHA re-authentication, and login requests. Accordingly, the collection pipeline can be continuously maintained even during long-term execution without user intervention. Ultimately, step (S310) is a step that provides a preliminary foundation for subsequent steps (S320) or (S350) to operate normally, and can be described as a foundation step that can be repeatedly called as needed during the execution of subsequent steps.
[0047] Next, the automatic determination step for the eligibility of court document collection (S320) is a step for automatically determining whether a specific case can actually be collected and in what way it should be processed. The possibility of obtaining court documents is not determined solely by the existence of a case number, but is influenced by the case type, case code, whether there are restrictions under regulations, whether internal portal code conversion is possible, and suitability of the registration system for each court. Therefore, in step (S320), the case code can first be extracted from the case number string. For example, if the case number is "2023Gohap123", the year can be separated into "2023", the case code into "Gohap", and the serial number into "123", and the extraction of the case code can be performed using a Korean character area identification method using regular expressions. Subsequently, the collection server (10) can perform a first determination by comparing the extracted case code with a list of excluded case codes based on legal regulations or operating rules. For example, case types excluded from electronic provision or copy application under the regulations, such as domestic cases, juvenile protection cases, victim protection order cases, victim child protection order cases, or prostitution-related protection cases, may be automatically excluded at this stage. Next, the collection server (10) can determine whether conversion is possible by comparing it with pre-registered case code-portal internal code mapping data. That is, since actual search or application API calls are possible only if the case code can be converted into the code required by the portal internally, if there is no mapping information, the case can be recorded as an item that cannot be automatically collected. In this case, the reason for failure is stored in the database, and a request for manual verification can be made to the person in charge through a messaging service. Meanwhile, if verification is passed, the case code is converted into the portal internal code, and the serial number can be normalized to a predetermined number of digits according to the portal's required format. In addition, since there may be a difference between the general court name recognized by the user and the registered name / registered code required by the portal or copy application system, double mapping can also be performed on the court name.As such, this step (S320) is a gateway step that reviews data consistency, legal permissibility, and system processing feasibility in a multi-layered manner before performing the actual collection process, and serves as a criterion for determining whether to perform payment or application in the subsequent step (S330) or to branch to separate manual processing.
[0048] Next, the API-RPA hybrid payment and automatic recovery step (S330) is a step in which, when a judgment is subject to paid viewing or paid provision, electronic means and user interface automation means are combined to create a state where it can be acquired. Although searching for some judgments on the court portal may be possible via API, actual purchase and downloading may require operations at the user interface level, such as loading into a shopping cart, card payment, and activating a secure payment module. Therefore, in step (S330), the number of unpaid items in the shopping cart can be checked first through an API call. If there are no unpaid items, the payment process can be skipped and the process can proceed to the next step; however, if unpaid items exist, the RPA tool can launch a browser to perform actions similar to those of an actual user. For example, the RPA tool can sequentially execute actions such as accessing the payment page, selecting the payment button, selecting a card company, entering card information, clicking the authentication button, and processing a secure payment module like an ISP. In this context, the term "hybrid" implies that while the verification of unpaid items and status inquiries are performed via API, the actual payment, which involves the intervention of a security program, is performed via UI operation. Furthermore, this step includes a payment automatic recovery function. That is, when a payment failure is detected, the collection server (10) can forcibly terminate security payment-related processes such as ISP, and if necessary, reinstall the security module and retry the payment. If this retry also fails, the request for manual intervention can be made by escalating to a person in charge via a messaging service such as Slack. When the payment is successfully completed, the browser window and payment-related processes can be automatically cleaned up to initialize the environment so as not to affect the next execution.According to the implementation example, step (S330) may be extended to include not only card payment but also a procedure to assist with fee payment after a copy request. In this case, the process may also include parsing the fee notification email to generate data for an account transfer request and retrieving the final PDF upon receiving a notification of completion of provision after payment. Ultimately, step (S330) is a step for securing actual acquisition rights for cases that received an eligibility determination in step (S320), and is closely linked with the authentication maintenance function of step (S310) so that if the session expires during payment, the authentication procedure is performed again and payment can be resumed.
[0049] Subsequently, the iterative restoration step (S340) is a step for progressively restoring the omitted parts through iterative re-searching and fragment combination when the entire judgment is not provided and only a preview text with parts omitted is provided. This is one of the important technical features of the present invention and serves to extend incomplete legal text beyond simple collection to a level where it can be substantially utilized. Specifically, the collection server (10) can first separate the text into multiple fragments based on an ellipsis, for example, "...", from the collected preview text. Each fragment becomes a text fragment in the form of an independent partial sentence or paragraph, and the collection server (10) can determine whether the beginning and end of each fragment are complete sentences. At this time, when determining completion, it is desirable to consider not only the presence of simple punctuation marks but also the termination patterns frequently used in legal documents. For example, "does.", "not guilty.", "dismissal.", "dismisses.", "rejects." The end_flag can be determined by checking whether a unique ending format of a judgment, such as the back, exists at the end of the fragment, and if the beginning of the fragment is abnormally cut off or starts without a particle or ending, the beginning part is determined to be omitted and the start_flag can be set to False. If it is determined that the beginning and end of all fragments are complete, the collection server (10) can consider this as a single complete text and proceed to the next step. However, if there are incomplete fragments, the text of a specific proportion area of the fragment, such as the front 1 / 3 or the back 1 / 3 area, can be extracted and morphological analysis can be performed to derive an extended keyword. In one embodiment, the morphological analysis can be performed using the Okt (Open Korean Text) morphological analyzer of KoNLPy, the input value is the text of a specific proportion area of the incomplete fragment, and the output value is a list of morphemes tagged with parts of speech.Afterward, only the part of speech of a noun is extracted, and stop words and duplicate words are removed, the first proper noun can be selected as an expanded keyword. The expanded keyword derived in this way is used as a search term to solve the problem of "what word should be re-searched to restore the omitted part." In other words, since it is difficult to reliably extract meaningful legal keywords using only simple string truncation or regular expressions, a morphological analysis model is utilized to derive actual re-search keywords from Korean legal sentences. Subsequently, the collection server (10) can re-search the original source or the same portal using the expanded keyword and obtain a new text fragment containing a longer context. Once the new fragment is obtained, the collection server (10) can determine whether there is an overlapping area of a predetermined length or more between the existing fragment and the new fragment; in one embodiment, the possibility of combination can be determined based on an overlapping string of 30 characters or more. The combination pattern can be applied by classifying it into multiple patterns such as identical (A=B), inclusion (BinA, AinB), forward concatenation (AB), and reverse concatenation (BA). For example, if the end of the existing fragment A and the beginning of the new fragment B overlap by 30 characters or more, they can be combined into an AB pattern, and conversely, if the end of B and the beginning of A overlap, they can be combined into a BA pattern. After combining, the collection server (10) can re-examine whether it is complete by determining the start_flag and end_flag again, and if it is still determined to be incomplete, it can repeat the extraction of expanded keywords based on morphological analysis, re-search, and fragment combination. Therefore, step (S340) is not a one-time post-processing step, but a cyclic step that is performed repeatedly as long as the possibility of restoration exists, and through this repeated execution, a much more complete legal text than the initial preview text can be obtained. Also, step (S340) is connected to step (S310) and step (S320).That is, if the session expires during the re-search process, step (S310) may be called again, and if a newly discovered event or document becomes a target for additional collection, the eligibility determination logic of step (S320) may be reapplied.
[0050] Finally, the data migration step (S350) is a step for transferring the judgment data collected, settled, restored, and refined in the preceding steps into an integrated service data structure, resolving duplication between multiple sources, and determining the final service availability state. Since judgments can be obtained through multiple collection channels, such as internet viewing paths, copy request paths, preview restoration paths, and public portal collection paths, there is a high probability that the same judgment will be stored redundantly in different sources. Accordingly, the collection server (10) can first determine whether the same judgment has been collected from multiple collection channels. In one embodiment, the determination of identity can be performed based on court and case_number_split, which is a method of determining that it is the same judgment if the combination of the court name and the normalized case number is identical. Subsequently, the collection server (10) can determine the source priority by comparing the source_type of the existing data with the source_type of the new data. For example, a source with higher fidelity to the original text, metadata completeness, structuring capability, or service quality may have a higher priority, while a source of lower quality or containing only limited information may have a lower priority. Based on the priority determination result, INSERT or UPDATE operations can be performed to disable the serviceable status of existing data or replace it with new data. In other words, even if identical judgments already exist, if the quality of the new collection is higher, the new case can be promoted to the primary service status, and the existing case can be disabled. Conversely, if the existing data originates from a higher-level source, the new case can be stored only as secondary data or excluded from migration. Additionally, reference IDs managed in the source table, such as case_id, can be updated with the new ID in the integrated table, ensuring subsequent referential integrity through this process. Subsequently, status fields such as analyzed and serviceable are updated, and search indexes or service caches can be updated as necessary.Ultimately, the final step (S350) is a final harmonization step that organizes heterogeneous data collected from multiple sources into a single integrated legal data asset, and serves to complete the results of the preceding steps so that they can be utilized for actual search, viewing, and service provision.
[0051] In summary, the method for collecting judgment documents according to one embodiment has a structure in which the steps of forming an authenticable state (S310), selecting actual collection possibilities and processing paths (S320), automatically securing paid acquisition rights (S330), repeatedly supplementing incomplete legal text (S340), and organizing results from multiple sources into integrated service data (S350) are organically linked. In particular, the present invention has a dynamic processing structure beyond a linear flowchart, in that step (S310) can be repeatedly called even during the execution of steps (S320) to (S340), step (S340) internally forms multiple re-search loops, and step (S350) absorbs the outputs of all preceding steps into a final data structure. Accordingly, the present invention can realize stable and serviceable collection of judgment document data by going beyond simple collection automation and comprehensively resolving all real-world problems such as authentication failures, payment failures, incomplete disclosure, and duplication of multiple sources.
[0052] Figure 4 is a diagram illustrating the function of the AI model in the iterative restoration step of Figure 3.
[0053] Referring to FIG. 4, in the iterative restoration step (S340), the AI model (11) may be implemented as a morphological analysis model for deriving an extended keyword valid for restoring an omitted section from an incomplete legal text. More specifically, the AI model (11) may receive at least one of an incomplete fragment front region (410) and an incomplete fragment back region (420) as input. Here, the incomplete fragment front region (410) may be a portion of a text fragment located before an ellipsis, and the incomplete fragment back region (420) may be a portion of a text fragment located after an ellipsis. In one embodiment, the incomplete fragment front region (410) may be set as the front 1 / 3 region of the incomplete fragment, and the incomplete fragment back region (420) may be set as the back 1 / 3 region of the incomplete fragment. However, the ratios are not limited thereto and may be set to various values depending on the length, sentence structure, or search precision requirements of the text to be restored.
[0054] For example, the entire text of an incomplete piece If so, the incomplete piece front region (410) can be defined as in the following mathematical formula 1.
[0055]
[0056] In addition, the area behind the incomplete piece (420) can be defined as shown in Equation 2 below.
[0057]
[0058] That is, the AI model (11) can selectively or together receive at least one of the preceding context and the succeeding context of the omitted section and analyze the key semantic words included in the corresponding context.
[0059] The AI model (11) can perform morphological analysis on the input incomplete fragment front region (410) or incomplete fragment back region (420) and, as a result, generate a morpheme list (330) tagged with parts of speech. In one embodiment, the AI model (11) may include the Okt (Open Korean Text) morpheme analyzer of KoNLPy. The AI model (11) can output a morpheme list (330) by breaking down the input text into morpheme units and assigning a part of speech tag corresponding to each morpheme. For example, the morpheme list (330) can be expressed as shown in Equation 3 below.
[0060]
[0061] Here, is the i-th morpheme, and may be a part-of-speech tag corresponding to the i-th morpheme. Therefore, the morpheme list (330) may not be a simple sequence of strings, but may be structured data in which each morpheme is linked with corresponding part-of-speech information.
[0062] The collection server (10) can extract noun candidates based on the above morpheme list (330). For example, a set of candidate nouns can be generated by extracting only morphemes whose part-of-speech tag is Noun, which can be defined as in Equation 4 below.
[0063]
[0064] Subsequently, the collection server (10) can remove words included in the pre-set stopword set and duplicate words from the above candidate noun set. Then, among the remaining candidates, the first proper noun in the order of appearance can be selected as an extended keyword. For example, the extended keyword q can be defined as shown in Equation 5 below.
[0065]
[0066] Here, S may be a set of stop words. In this way, the direct output of the AI model (11) may be a list of morphemes (330), and an extended keyword used for subsequent re-search may be derived based on the list of morphemes (330).
[0067] According to the above configuration, the AI model (11) can identify key nouns in Korean legal sentences that are difficult to extract using only simple string splitting or regular expression processing. In particular, since Korean legal sentences often involve complex combinations of particles, endings, conjunctions, and modifiers, it is difficult to reliably extract keywords that are practically valid for re-searching based on simple strings. On the other hand, the AI model (11) can more accurately derive search terms suitable for restoring omitted sections by analyzing the context of the incomplete fragment front region (410) and the incomplete fragment back region (420) in morpheme units and providing part-of-speech information. Accordingly, the AI model (11) can function as a keyword selection means to solve the problem of “which word should be used for re-searching to restore the omitted part.”
[0068] Meanwhile, the AI model (11) is not limited to simply calling an existing morphological analyzer, but can be trained separately to improve keyword selection performance for restoring incomplete legal text. Below, an example of a training method for the AI model (11) is described.
[0069] In one embodiment, a collection server (10) or a separate learning server may generate an incomplete learning sample that includes an artificially omitted section from the complete original text of the judgment. For example, the complete original text of the judgment Let it be, and any continuous interval among them If is set as the omission target, the actual omission interval Y can be defined as shown in Equation 6 below.
[0070]
[0071] And, by replacing the above omitted section Y with an ellipsis symbol, the incomplete learning sample It can generate, which can be expressed as in mathematical formula 7 below.
[0072]
[0073] Subsequently, the above incomplete learning sample The context before and after the ellipsis can be extracted from to generate an incomplete fragment front region (410) and an incomplete fragment back region (420). Then, the AI model (11) can perform morphological analysis on the front and back regions to extract multiple candidate nouns. Each candidate noun For , perform a re-search to restore candidate text You can obtain.
[0074] At this time, in order to quantify the restoration contribution of each candidate noun, the restoration utility value It can be calculated. In one embodiment, the restoration utility value duplicate points , completion score , semantic similarity score and position weights It can be defined as in the following mathematical formula 8 by combining them.
[0075]
[0076] Here, and can be a weight that adjusts the contribution of each element. The above duplicate score is restoration candidate text The degree of overlap between and the actual omitted interval Y can be represented, and can be defined, for example, as in Equation 9 below.
[0077]
[0078] Here, is restoration candidate text It may be the longest duplicate string length between and the actual omitted section Y, or the sum of consecutive duplicate strings of a predetermined length or longer. In one embodiment, the predetermined length may be 30 characters.
[0079] The above completion score It can be determined based on whether the combined text after restoration satisfies legal document closing patterns. For example, it can be set to have a value of 1 if the text after restoration satisfies closing patterns such as "does," "not guilty," "acquittal," or "dismisses," and a value of 0 otherwise. In addition, the above semantic similarity score is restoration candidate text It can be calculated as the cosine similarity between the sentence embeddings of the actual omitted interval Y and, which can be expressed as Equation 10 below.
[0080]
[0081] Here, can be a sentence encoder. The above position weights can reflect the degree to which a candidate noun is close to the ellipsis boundary, and can be defined, for example, as in Equation 11 below.
[0082]
[0083] Here, is the distance between the candidate noun and the ellipsis boundary, and can be the total length of the input area.
[0084] The collection server (10) has the above restoration utility value Target distribution for candidate nouns using It can generate. For example, the above target distribution It can be calculated using the softmax function as shown in Equation 12 below.
[0085]
[0086] Here, can be a temperature parameter. Meanwhile, the keyword selection model for training predicts the score for each candidate noun. Calculate and predict distribution It can generate. For example, the above prediction distribution It can be defined as shown in the following mathematical formula 13.
[0087]
[0088] In addition, the above predicted score is the context embedding of the incomplete fragment front region (410) , context embedding of the back region (420) of the incomplete fragment , and embeddings of candidate nouns It can be calculated by combining, for example, as defined in mathematical formula 14 below.
[0089]
[0090] Here, and can be a learnable parameter.
[0091] The AI model (11) is the above target distribution and predicted distribution It can be learned to minimize the difference between. In one embodiment, the loss function It can be defined as shown in the following mathematical formula 15.
[0092]
[0093] Here, ε₀ may be the index of the candidate noun with the highest restoration utility value, and the first term may be the cross-entropy loss between the target distribution and the prediction distribution. Additionally, the second term may be the rank margin loss that ensures the optimal candidate noun has a score higher than other candidates by a certain margin, and the third term may be a normalization term for the model parameters.
[0094] According to the above learning method, the AI model (11) can be trained to go beyond simply classifying morphemes by part of speech and to prioritize the selection of nouns with a high actual restoration success rate. In particular, by constructing training data in a self-supervised manner by generating artificial omission sections from complete judgments and automatically generating teacher signals based on the overlap, completeness, semantic similarity, and boundary proximity between the actual re-search results and the omission sections, restoration performance can be directly optimized without separate manual labeling. Accordingly, the AI model (11) can be utilized as a technical means to improve re-search accuracy and omission section restoration efficiency in the iterative restoration step (S340).
[0095] FIG. 5 is a diagram illustrating the hardware configuration of a collection server according to one embodiment of the present invention.
[0096] Referring to FIG. 5, a collection server (10) according to one embodiment of the present invention may be implemented as an electronic device for performing functions of collecting, authenticating, restoring, storing, and providing data of judgments. The collection server (10) may include a processor (110), memory (120), a transmitting / receiving device (130), an input interface device (140), an output interface device (150), a storage device (160), and a bus (170), and an AI model (11) may perform functions of morphological analysis, derivation of expanded keywords, re-search assistance, and restoration judgment in conjunction with the above components.
[0097] The processor (110) is a component that controls the overall operation of the collection server (10) and can execute judgment collection logic, automatic authentication logic, payment automation logic, iterative restoration logic, and data migration logic. In particular, the processor (110) can execute program instructions stored in memory (120) to perform access to external legal data sources, authentication token processing, case code verification, incomplete fragment analysis, and integrated data generation. Memory (120) may include ROM and RAM and can store operating programs, authentication information, temporary collected data, text to be restored, morpheme lists, and execution code or parameters of the AI model (11).
[0098] The transmitting and receiving device (130) is a component for transmitting and receiving data with an external data server, a user terminal, or other linked server. Through this, the collection server (10) can transmit and receive the original text of the judgment, metadata, authentication response, payment status information, and user request information. The input interface device (140) is a component for receiving a case number, search keyword, authentication information, or control command from an administrator or user, and the output interface device (150) may be a component for providing collection results, restoration results, status information, or warning messages to the outside.
[0099] The storage device (160) is a component for long-term storage of collected original judgments, structured data, logs, duplicate determination results, migration results, and training data. The bus (170) can provide a data transfer path between the processor (110), memory (120), transmission / reception device (130), input interface device (140), output interface device (150), and storage device (160). Accordingly, the collection server (10) of FIG. 5 can be understood as an integrated electronic device in which each hardware component is organically connected to reliably perform the judgment collection method of the present invention.
[0100] Although embodiments according to the technical concept of the present invention have been described above with reference to the attached drawings, those skilled in the art will understand that the present invention may be implemented in other specific forms without changing its technical concept or essential features. The embodiments described above should be understood as illustrative in all respects and not restrictive. Explanation of the symbols
[0102] 10: Collection Server, 11: AI Model, 12: Database, 20: Data Server, 30: User Terminal, 110: Processor, 120: Memory, 130: Transmitter / Receiver, 140: Input Interface Device, 150: Output Interface Device, 160: Storage Device, 170: Bus, 210: Judgment Collection Module, 220: Automatic Authentication Module, 230: Hybrid Payment Automation Module, 240: Incomplete Legal Text Restoration Engine, 250: Multi-layered Verification and Court Code Dual Mapping System, 260: Judgment Structure Automatic Decomposition Engine, 330: Morphological List, 410: Front Region of Incomplete Fragment, 420: Back Region of Incomplete Fragment
Claims
Claim 1 In an AI inference-based judgment collection system, a data server that provides legal data including judgments; and a user terminal that receives a user's request to collect judgments and outputs the judgment collection results; A collection server is connected to communicate with the data server and the user terminal, and includes an AI model that analyzes request information transmitted from the user terminal and previously secured data to derive a direction for the collection or processing of the judgment, and a database that stores the original text of the judgment, metadata, collection history, user request history, AI inference results, duplicate determination results, and refined structural data; wherein the collection server comprises: a judgment collection module that obtains at least one of the original text of the judgment, preview text, metadata, attachments, or download paths from the data server; an automatic authentication module that extracts an authentication token to allow the judgment collection module to access an external portal, sets a cookie, performs CAPTCHA verification and login processing to form an access session, and automatically restores the access session upon session expiration; a hybrid payment automation module that, when paid viewing or fee payment is required, performs payment by combining API-based status inquiry and user interface automation according to the request of the judgment collection module, and performs a recovery operation in the event of payment failure; and extracts a case code from the case number included in the request information, verifies whether it is subject to exclusion and whether internal code conversion is possible, and, upon passing the verification, the court name and case code on the portal A multi-layered verification and court code dual mapping system that converts into an internal identification system and provides it to the above-mentioned judgment collection module;A judgment collection system comprising: an incomplete legal text restoration engine that, when data acquired by the judgment collection module is preview text or partially omitted legal text, separates fragments based on ellipsis symbols, determines whether each fragment is complete, derives expanded keywords from incomplete fragments and requests a re-search to the judgment collection module, and combines new fragments acquired through re-search by the judgment collection module with existing fragments; and an automatic judgment structure decomposition engine that extracts at least some of case expressions, case names, ruling matters, summary of judgment, referenced provisions, referenced precedents, orders, claims, reasons, and judge information from the original text acquired by the judgment collection module or the text restored by the incomplete legal text restoration engine to generate structured data, and provides the generated structured data to be stored in the database. Claim 2 delete Claim 3 A judgment collection system according to claim 1, wherein the AI model derives a collection strategy including collection priority, collection target, and restoration target based on previously secured data stored in the database and request information input from the user terminal, and provides this strategy to the judgment collection module; the judgment collection module acquires a judgment from the data server using the collection strategy, an access session formed by the automatic authentication module, and a portal internal identification system provided by the multi-layer verification and court code dual mapping system; the judgment collection module calls the hybrid payment automation module if paid viewing or fee payment is required during the acquisition process, and calls the incomplete legal text restoration engine if the acquired data is preview text or partially omitted legal text; the judgment structure automatic decomposition engine structures the text provided by the judgment collection module or the incomplete legal text restoration engine; the database stores the structured data; and the user terminal is configured to receive the structured data stored in the database or the judgment collection results. Claim 4 A judgment collection system according to claim 1, wherein the AI model receives text of the front or back area of the incomplete fragment as input, performs morphological analysis, and outputs a list of morphemes tagged with parts of speech, and the incomplete legal text restoration engine is configured to extract nouns from the list of morphemes, remove stop words and duplicate words, and then select the first proper noun as an expanded keyword. Claim 5 A judgment collection system according to claim 4, wherein the input to the AI model includes text of the front 1 / 3 area or the back 1 / 3 area of an incomplete fragment, and the AI model is configured to include the Okt morphological analyzer of KoNLPy. Claim 6 In Paragraph 4, the above AI model is the complete original text of the judgment From any continuous interval Set as the target for omission, and define the actual omitted interval Y as in Mathematical Formula 1 below, and <Mathical Formula 1> ,Incomplete learning sample with the above omitted section Y replaced by an ellipsis Generate as shown in Mathematical Formula 2 below, and <Mathical Formula 2> A judgment collection system that obtains restoration candidate texts by performing a re-search on each of a plurality of candidate nouns extracted from the front and rear regions of the above-mentioned incomplete learning sample, calculates restoration utility values for each candidate based on the overlap, completeness, semantic similarity, and boundary proximity between the restoration candidate texts and the actual omitted sections, and is trained to minimize the difference between the target distribution based on the restoration utility values and the model's prediction distribution. Claim 7 In Clause 6, the AI model is each of the above candidate nouns The above restoration candidate text for and the restoration utility value for each candidate based on the actual omitted section Y above Calculate according to the following mathematical formula 3, and <Mathematical Formula 3> , duplicate points in the above mathematical formula 3 is calculated according to the following mathematical formula 4, and <Mathematical Formula 4> , semantic similarity score is calculated according to the following mathematical formula 5, and <Mathematical Formula 5> position weights is calculated according to the following mathematical formula 6, and <Mathematical Formula 6> , above and is a weight, and the above is a score indicating the completeness of the combined text after restoration, and the above function is a sentence encoder, and the above is the distance between the candidate noun and the ellipsis boundary, and the above A judgment collection system, which is the total length of the input area. Claim 8 delete Claim 9 A method for collecting judgments based on AI inference performed on a server including an AI model, comprising: a multi-stage automatic authentication and session automatic recovery step of extracting an authentication token from the main page of a web portal provided by an external server, setting the authentication token in a cookie, performing CAPTCHA verification and login processing to form an access session for collecting judgments, and automatically recovering the access session by re-extracting the authentication token, setting the cookie, performing CAPTCHA verification and login processing if session expiration is detected during the execution of a subsequent step; an automatic court document collection eligibility determination step of extracting a case code from a case number based on the access session, verifying whether the case is subject to exclusion and whether internal portal code conversion is possible, and generating request parameters for subsequent collection or application by converting the court name and case code for collectible cases into an internal identification system; an API-RPA hybrid payment and automatic recovery step of performing payment by combining API-based status inquiry and user interface automation when paid viewing or fee payment is required as a result of performing a judgment search or application using the request parameters, performing a recovery action in case of payment failure, and acquiring the judgment or judgment-related data after payment is completed; and the judgment Alternatively, if the data related to the judgment is preview text or partially omitted legal text, it is separated into multiple fragments based on ellipsis symbols, the completion status of each fragment is determined, extended keywords are derived from incomplete fragments to perform a re-search, and new fragments and existing fragments are combined, wherein if session expiration is detected during the re-search, the above-mentioned multi-step automatic authentication and session automatic recovery steps are re-executed to continue the re-search, and if it remains incomplete even after combining, the above-mentioned extended keywords are derived, and the re-search and combination are repeated in an iterative restoration step;A method for collecting judgments, comprising a data migration step of determining identity by comparing a judgment obtained in the API-RPA hybrid payment and automatic recovery step or a judgment restored in the iterative restoration step with other judgments obtained from multiple collection channels, changing the service availability status or replacing it with new data according to source priority, and updating the reference ID of the original table with the new ID of the integrated table. Claim 10 A method for collecting judgments according to claim 9, wherein the iterative restoration step comprises the AI model receiving at least one text of the front 1 / 3 area and the back 1 / 3 area of the incomplete fragment as input, performing morphological analysis to generate a list of morphemes tagged with parts of speech, extracting nouns from the list of morphemes, removing stop words and duplicate words to select the first proper noun as an expanded keyword, and performing a re-search using the selected expanded keyword.