Ultra-large ai-based technical document risk factor detection system and method
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- BIGWAVE AI CO LTD
- Filing Date
- 2024-11-29
- Publication Date
- 2026-06-04
Smart Images

Figure KR2024019270_04062026_PF_FP_ABST
Abstract
Description
Ultra-large AI-based technical document risk factor detection system and method
[0001] The present invention relates to a system and method for detecting risk factors in technical documents based on a massive AI that automatically detects risk factors in parts where design changes are expected. In particular, it relates to a system and method for detecting risk factors in technical documents based on a massive AI that enables question-and-answer interaction through generative AI even for non-experts in the relevant field and allows for the acquisition of reliable information based on technical documents, thereby enhancing the services of a large language model (LLM)-based design document analysis support solution for the construction field and enabling the establishment of advanced strategies.
[0002] The present invention is a technology developed by Big Wave AI Co., Ltd. (project implementing agency name) through "Research Project Name: Demonstration of Ultra-large AI-based Technology Document Risk Factor Detection Solution (Project No.: SW-Voucher-06, Research Period: 2024-07-01 ~ 2024-11-30)" as part of the Digital Innovation Enterprise Global Growth Voucher Support Project (Research Project Name) of the Ministry of Science and ICT (Ministry Name) of the Republic of Korea and the Korea Information and Communication Industry Promotion Agency (Research Management Agency).
[0003]
[0004] Recently, in natural language processing research, the approach utilizing LLMs has garnered significant attention for demonstrating overwhelmingly superior performance in various fields, including document processing and summarization, compared to the use of existing small-scale language models. It is known that to maximize the performance of large-scale language models, reinforcement learning is required in addition to the conventional method of training pre-trained models to fit a given task using fine-tuning. The reinforcement learning performed in this context is primarily Reinforcement Learning from Human Feedback (RLHF), which may require a process in which humans directly evaluate the quality of the results generated by the fine-tuned large-scale language model. Furthermore, reinforcement learning in RLHF is mainly carried out through Proximal Policy Optimization (PPO), an online method. PPO may require a reward model capable of evaluating in real-time the quality of the examples generated by the large-scale language model being trained.
[0005] For reference, in addition to the large-scale language model, a reward model must be additionally trained on the given task. There is a problem in that performing reinforcement learning on a large-scale language model using the Human Feedback-Based Reinforcement Learning (RLHF) method requires a significant amount of effort and hardware resources because humans must directly evaluate the results generated by the model and multiple models must be trained on the same task.
[0006] In addition, service industries providing consultation services offer consultation chatbots using LLM-based artificial intelligence with their own proprietary and developed content. Consultation chatbots utilizing LLM AI models are employed for consultation services that use AI RLHF, such as ChatGPT4 and Gemini. Consultation chatbots based on AI models utilize natural language processing and machine learning technologies to provide human-like conversations, thereby enabling high operational efficiency and the collection of various data.
[0007] Meanwhile, efficient information management refers to establishing and operating a system that enables the sending, dissemination, storage, retrieval, and retrieval of documents so that the necessary people can access them at the necessary time and place; however, most current project management systems prioritize only information storage, resulting in low utilization and losses in project execution due to the inability to secure necessary information when problems arise during project implementation.
[0008] In other words, while the construction industry utilizes various business systems such as ERP (Enterprise Resource Planning) and groupware for document management, there are documents—including product and process technical documents—that are difficult for non-experts to understand. Consequently, the integration of information retrieval systems and document searching become challenging due to the diverse business systems and big data.
[0009] An example of a technology for solving these problems is disclosed in the following patent documents 1 to 3, etc.
[0010] For example, Patent Document 1 (Republic of Korea Registered Patent Publication No. 10-2074578, registered on January 31, 2020) describes an official document processing module that processes and stores data from a sender / receiver box where official documents are sent and received; a parsing module that parses the content and history of negative words and keywords included in the data stored through the official document processing module and stores it as structured data; an ID generation module that assigns and manages morpheme-based IDs for the utilization of data constructed through the parsing module; a correlation analysis module that analyzes correlations based on the sender / receiver, destination, positive / negative status, and design-related status through the data constructed through the official document processing module, parsing module, and ID generation module, and identifies the client's tendencies through the derived correlations; a utilization data configuration module that outputs the analysis results in the form of a report for the utilization of the analysis results analyzed through the correlation analysis module; and the official document processing module, parsing module, ID generation module, and correlation A text mining-based construction document analysis system is disclosed, comprising a management server that transmits, receives, and manages information in conjunction with an analysis module and a utilization data configuration module.
[0011] In addition, Patent Document 2 (Republic of Korea Registered Patent Publication No. 10-2699424, registered on August 22, 2024) discloses a system for providing a consultation chatbot that receives a question text from an external electronic device connected to a server, extracts a prompt to be input into an artificial intelligence model based on a large-scale language model from the question text, generates a response message based on content stored in a database by inputting the prompt into the artificial intelligence model based on the large-scale language model, provides the response message to the external electronic device through the virtual space of the server, counts the amount or number of times the response message is provided, and collects and analyzes the communication log of the external electronic device based on identifying that at least one of the accumulated amount or number of times the response message is provided exceeds a threshold, and based on the analysis result, identifies whether the external electronic device is a member, continues to provide the response message if it is identified as a member, and stops providing the response message if it is identified as a non-member.
[0012] Meanwhile, Patent Document 3 (Republic of Korea Registered Patent Publication No. 10-2647511, registered on March 11, 2024) discloses a reinforcement learning method for a large-scale language model that includes the steps of acquiring input data for a language model and outputting result data for a natural language processing task using the language model, wherein the language model sets a reward function based on a reward for original training data and a reward for augmented training data, and the language model is a reinforcement learning language model such that the reward obtained according to the reward function is maximized during the process of the language model performing the natural language processing task, the reward for the original training data is a fixed reward, and the reward for the augmented training data is an automated dynamic reward.
[0013] Patent Document 1, as described above, discloses a technology that reduces the time required for bid review by IDing various standards and enables quick response in the event of a problem by IDing standard documents such as contracts; Patent Document 2 discloses a technology that enables accurate and rapid consultation by solving the problem of providing incorrect information to users (customers) due to responses unrelated to the content; and Patent Document 3 discloses a technology for reinforcement learning of LLM, but does not disclose generative AI such as ChatGPT.
[0014] Meanwhile, existing generative AI models such as ChatGPT face the problem of hallucination, where they respond as if incorrect information were factual, making it difficult to acquire reliable information. Additionally, due to the nature of the construction industry, there was a problem in that commercial SaaS (Software as a Service) services could not be applied because of security issues regarding the risk of information leakage from internal corporate documents, such as technical and design documents, which could lead to the external leakage of internal information.
[0015] In addition, there was a problem in that it was difficult for local government agencies and small and medium-sized enterprises to adopt LLM models for self-built services due to the high cost burden in terms of computing resources and time when proceeding with pre-training.
[0016] The objective of the present invention is to solve the problems described above by providing a system and method for detecting risk factors in ultra-large AI-based technical documents that enables question-and-answering through generative AI and allows reliable information based on technical documents to be obtained even by non-experts in the relevant field.
[0017] Another objective of the present invention is to provide a system and method for detecting risk factors in ultra-large AI-based technical documents that can ensure the reliability of the LLM model by utilizing search augmentation techniques of RAG (Retrieval-Augmented Generation) together with a Large Language Model (LM) to generate answers by referencing a vast amount of documents held by an enterprise, and by designing the system not to provide answers when there is no reference material available.
[0018] Another objective of the present invention is to provide a super-large AI-based technical document risk factor detection system and method capable of efficiently detecting design change risks within new design documents by managing design document review know-how in a knowledge base through design change history data held by an organization and applying it to a RAG-based LLM model.
[0019] To achieve the above objective, the ultra-large AI-based technical document risk factor detection system according to the present invention is a system for detecting risk factors in technical documents based on ultra-large AI, and is characterized by comprising: a data collection module for collecting analog and digital documents held within an organization; a text extraction module for automatically extracting text from the analog and digital documents collected by the data collection module through OCR and / or Text Block recognition; a data preprocessing pipeline construction module for augmented search that divides the text extracted by the text extraction module into units recognizable by generative AI and vectorizes it; a design change detection module for automatically detecting parts within a design document where design changes are expected based on design change history data collected by the data collection module; and a management server that transmits, receives, and manages information in conjunction with the data collection module, the text extraction module, the data preprocessing pipeline construction module, and the design change detection module.
[0020] In addition, the technical document risk factor detection system according to the present invention is characterized by further including a storage module that stores meta-information extracted from the text extraction module and vector information from the preprocessing pipeline construction module.
[0021] In addition, in the technical document risk factor detection system according to the present invention, the design change detection module is characterized by detecting design change documents according to the project name, construction period, construction amount, type of construction, and options.
[0022] In addition, in the technical document risk factor detection system according to the present invention, the design change detection module detects and outputs a history inquiry of a similar construction design change similar to the construction information entered by the user, and the similar construction design change history information includes information on a construction overview, design change items, reasons for design changes, change costs, and reference document information, wherein the construction overview includes a construction name, construction period, construction amount, type of construction, and contractor, and the reference document information is characterized as being the original image of the design change.
[0023] Furthermore, the technical document risk factor detection system according to the present invention further includes a question-response module that answers a question entered by a user with information regarding a design change detected by the design change detection module, wherein the question-response module provides a reference document for the response, and the reference document is characterized as being an original document image of the design change document information used to generate the response.
[0024] In addition, to achieve the above objective, the method for detecting risk factors in technical documents according to the present invention is a method for detecting risk factors in technical documents based on a super-large AI, and is characterized by comprising: (a) a data collection step of collecting analog documents and digital documents held within an organization; (b) a text extraction step of automatically extracting text from the analog documents and digital documents collected in step (a) through OCR and / or Text Block recognition; (c) a step of constructing a data preprocessing pipeline for augmented search by dividing the text extracted in step (c) into units recognizable by generative AI and vectorizing it; (d) a design change detection step of automatically detecting parts within a design document where design changes are expected based on design change history data collected in step (a); and (e) a step of responding to a question entered by a user with information regarding the design changes detected in step (d).
[0025] As described above, according to the ultra-large AI-based technology document risk factor detection system and method of the present invention, the effect of significantly improving work efficiency through automated document processing and information retrieval functions is obtained.
[0026] Furthermore, according to the ultra-large AI-based technical document risk factor detection system and method of the present invention, the time required for repetitive document tasks, document reading, searching, etc., is reduced, thereby achieving the effect of improving the economic efficiency and competitiveness of the design and review process in the construction safety field.
[0027] In addition, according to the ultra-large AI-based technical document risk factor detection system and method of the present invention, the effect of enabling agile market response and innovation by shortening the time for technical document review and decision-making is also obtained.
[0028] FIG. 1 is a diagram illustrating the concept of general text mining applied to the present invention,
[0029] FIG. 2 is a block diagram of a technical document risk factor detection system according to the present invention,
[0030] FIG. 3 is a diagram showing an example of a RAG augmented search pipeline construction module illustrated in FIG. 2.
[0031] FIG. 4 is a flowchart illustrating a method for detecting risk factors in ultra-large AI-based technical documents according to the present invention,
[0032] FIGS. 5 to 22 are drawings illustrating an example of the operation of a super-large AI-based technology document risk factor detection system according to the present invention.
[0033] The above and other objects and novel features of the present invention will become more apparent from the description in this specification and the accompanying drawings.
[0034] The size and thickness of each component shown in the description and drawings of the present invention are depicted arbitrarily for convenience of explanation, and therefore the present invention is not necessarily limited to what is illustrated.
[0035] Meanwhile, in the description of the invention, when it is stated that a certain part "includes" a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components.
[0036] The terms “part,” “module,” or “part” as used herein perform at least one function or operation and may be implemented as hardware or software consisting of mechanical or electrical / electronic configurations, or as a combination of hardware and software; and a plurality of “parts,” “modules,” or a plurality of “parts” may be integrated into at least one module and implemented by at least one processor, except for the “parts,” “modules,” or “parts” that need to be implemented in specific hardware.
[0037] Text mining applicable to the present invention is a technology for finding valuable and meaningful information from unstructured text, as illustrated in FIG. 1. It utilizes natural language processing technology that integrates linguistics, statistics, machine learning, etc. to structure semi-structured / unstructured text data, extract features, and derive meaningful information from the extracted features.
[0038] Such text mining applies word frequency analysis, cluster analysis, topic modeling, sentiment analysis, and association analysis techniques, and recently it is widely used in risk management, knowledge management, customer service, consumer patterns, and social network service analysis.
[0039] The aforementioned text mining process may include data collection for collecting data necessary for analysis, such as contract documents, through optical character recognition, PDF recognition, PPT recognition, API (Application Programmer Interface) / RPA (Robotic Process Automation) / web crawling; data preprocessing for processing the collected data into a suitable form by analyzing it, such as removing unnecessary phrases and restoring the original form; text analysis for deriving meaningful information through keyword extraction and document classification based on the processed data; and visualization for visually expressing the analysis results.
[0040] In the system and method for detecting risk factors in ultra-large AI-based technical documents according to the present invention, text is automatically extracted from analog and digital documents held within an organization through OCR, Text Block recognition, etc., and the extracted text is divided into units recognizable by generative AI and vectorized to construct a data preprocessing pipeline for augmented search. Furthermore, based on design change history data, parts within design documents where design changes are expected are automatically detected and corrective measures are provided. A system can be constructed that allows querying regarding content related to design change documents and verifying the results of the answers to the queries related to design change documents.
[0041] Hereinafter, an embodiment according to the present invention will be described with reference to the drawings.
[0042] FIG. 2 is a block diagram of a technical document risk factor detection system according to the present invention.
[0043] The technical document risk factor detection system according to the present invention is a system for detecting risk factors in technical documents based on a super-large AI, and as illustrated in FIG. 2, comprises: a data collection module (100) for collecting analog documents and digital documents held within an institution; a text extraction module (200) for automatically extracting text from analog documents and digital documents collected by the data collection module (100) through OCR and / or Text Block recognition; a data preprocessing pipeline construction module (300) for constructing a data preprocessing pipeline for augmented search by dividing the text extracted by the text extraction module (200) into units recognizable by a generative AI and vectorizing it; a design change detection module (400) for automatically detecting parts where design changes are expected within a design document based on design change history data collected by the data collection module (100); a storage module (500) for storing meta information extracted by the text extraction module (200) and vector information from the preprocessing pipeline construction module (300); and a question-answer module (600) for answering a question entered by a user with information on design changes detected by the design change detection module (400). It may include a management server (700) that transmits, receives, and manages information in conjunction with the above data collection module (100), text extraction module (200), data preprocessing pipeline construction module (300), design change detection module (400), storage module (500), and question response module (600).
[0044] In addition, the technical document risk factor detection system according to the present invention may further include an input module equipped with a keyboard, mouse, scanner, etc., for inputting information for data collection in the data collection module (100) or detection in the design change detection module (400), and an output module including a monitor, etc., for displaying detection information in the design change detection module (400) and response information in the question and answer module (600).
[0045] The above data collection module (100) collects necessary information from analog documents such as printed official documents held within the institution, HWP file documents, WORD file documents, PDF documents, or digital documents such as PPT, and from the collected data, document data regarding the construction name, construction period, construction cost, type of construction, design change history information, etc., can be collected as technical documents according to the present invention, for example, through OPEN API (Application Programmer Interface) / RPA (Robotic Process Automation) / web crawling.
[0046] The text extraction module (200) extracts text from documents collected by the data collection module (100), for example, regarding the construction name, construction period, construction amount, type of construction, and design change history information such as construction overview, design change items, reasons for design changes, and change costs, and may also extract information in the form of design drawings and tables. Additionally, the text information meta-information extracted by the text extraction module (200) may be stored in the storage module (500).
[0047] As shown in FIG. 3, the data preprocessing pipeline construction module (300) divides the text extracted from the text extraction module (200) into units that can be recognized by a generative AI and then vectorizes it, and this vectorized information may also be stored in the storage module (500). FIG. 3 is a diagram showing an example of the RAG augmented search pipeline construction module shown in FIG. 2.
[0048] FIG. 3 illustrates a document regarding the Daegu City Development Corporation as an example of a data preprocessing pipeline construction module (300). The data preprocessing pipeline construction module (300) can convert text extracted from a text extraction module (200) into a vector form by dividing it into units that can be recognized by a generative AI and embedding it, perform preliminary research on a generative AI-based text embedding model, design a structure for indexing original document metadata and embedding vectors, and store it as a vector DB for efficiently searching for document information related to a question in a question-answer module (600).
[0049] The design change detection module (400) can detect design change documents based on the project name, construction period, construction amount, type of construction, and selection options in the documents extracted and stored by the text extraction module (200), and detects and outputs a history lookup of similar construction design changes similar to the construction information entered by the user. The similar construction design change history information includes information on the project overview, design change items, reasons for design changes, change costs, and reference document information. The project overview may include the project name, construction period, construction amount, type of construction, and contractor. The reference document information may be visualized and output on a monitor, which is an output module, as an image of the original (document) of the design change.
[0050] In addition, the search category of the above-mentioned design change detection module (400) can realize context-aware search unlike conventional keyword-based search, and can reduce the time required by automating risk clause search and history provision.
[0051] The above storage module (500) may be composed of memory elements and may be constructed as a database. When developing a specific device, the configuration of such a database may be structured according to database construction theory, taking into account ease of access and search and efficiency.
[0052] The above question-response module (600) responds to a query entered by a user through an input module by using meta information extracted from the text extraction module (200) and vector information from the preprocessing pipeline construction module (300) according to the information stored in the storage module (500), and can provide a reference document for the response, and the reference document may be an original document image of the design change document information used to generate the response.
[0053] The above management server (700) may be configured to function as a type of web server or web service server. For example, the management server (700) displays extracted information on a webpage of an output module or receives necessary input data through a webpage. The webpage includes not only simple text, images, and multimedia, but also driving software for performing specific tasks, such as web applications. Additionally, the above management server (700) may be provided as a dedicated server for the technical document risk factor detection system according to the present invention, but is not limited thereto.
[0054] Next, the process of establishing a risk factor detection system for ultra-large AI-based technical documents according to the present invention will be explained with reference to Fig. 4.
[0055] Figure 4 is a flowchart illustrating a method for detecting risk factors in ultra-large AI-based technical documents according to the present invention.
[0056] The method for detecting risk factors in a technical document based on a super-large AI according to the present invention is a method for detecting risk factors in a technical document based on a super-large AI, and as illustrated in FIG. 4, a data collection module (100) collects analog documents and digital documents held within the institution (S10). In step S10, necessary information can be collected from analog documents such as printed official documents held within the institution, and digital documents such as HWP file documents, WORD file documents, PDF documents, or PPT.
[0057] Next, text is automatically extracted from the analog and digital documents collected in step S10 by the text extraction module (200) through OCR and / or Text Block recognition (S20). In step S20, the text extraction module (200) extracts text from the documents collected by the data collection module (100), for example, regarding the construction name, construction period, construction amount, type of construction, and design change history information such as construction overview, design change items, reasons for design changes, change costs, etc., and may also extract information in the form of design drawings and tables.
[0058] Additionally, the text extracted in step S20 is divided into units recognizable by the generative AI and vectorized to construct a data preprocessing pipeline for augmented search (S30). That is, in step S30, the text extracted from the text extraction module (200) is divided into units recognizable by the generative AI in the data preprocessing pipeline construction module (300), embedded, and converted into a vector form, and a preliminary study on a generative AI-based text embedding model can be performed, and a structure for indexing original document metadata and embedding vectors can be designed.
[0059] In addition, based on the design change history data collected in step S10, the design change detection module (400) automatically detects parts of the design document where a design change is expected (S40). That is, the design change detection module (400) can detect design change documents based on the construction name, construction period, construction cost, construction type, and selection options in the documents extracted and stored by the text extraction module (200), and can detect history inquiries of similar construction design changes similar to the construction information entered by the user.
[0060] Meanwhile, in the method for detecting risk factors in ultra-large AI-based technical documents according to the present invention, a response to a question entered by a user can be provided with information regarding design changes detected in step S40 (S50). The response in step S50 provides a reference document, and the reference document can be provided as an image of the original document of the design change document information used to generate the response.
[0061] Next, the operation of a super-large AI-based technical document risk factor detection system as described above will be explained with reference to FIGS. 5 to 22. FIGS. 5 to 22 are drawings illustrating an example of the operation of a super-large AI-based technical document risk factor detection system according to the present invention.
[0062] First, as illustrated in FIG. 5, the user can access the ultra-large AI-based technical document risk factor detection system according to the present invention through an input module. That is, the user can access the system according to the present invention by clicking the [Login] button after entering an account ID and password assigned by an administrator. The account ID and password can be set to be the same as, for example, an employee number.
[0063] When logged in, access to the design change history inquiry analysis system according to the present invention can be established, as illustrated in FIG. 6. In such a connection, the initial screen of the output module may be displayed in a state of inputting construction information for detecting design changes.
[0064] Meanwhile, in the ultra-large AI-based technology document risk factor detection system according to the present invention, file management can be performed by selecting a file management item as a RAG service as shown in FIG. 7, and then viewing currently uploaded design change-related documents as shown in FIG. 8. These files can be displayed by classifying them by filename, author, creation date, and upload status. Additionally, in the present invention, as shown in FIG. 9, a keyword can be entered and a search button can be clicked to view a list of files containing the keyword.
[0065] Meanwhile, as illustrated in FIG. 10, the user or administrator may select a file to download and download it, or delete an uploaded file. Additionally, as illustrated in FIG. 11 and FIG. 12, design change related documents may be additionally uploaded. When such files are uploaded, the text extraction module (200) may initiate data extraction.
[0066] Next, the process of operating design change detection in the design change detection module (400) according to the present invention will be described. When the Design Change Detection > Construction Information Input button is clicked in FIG. 13, as shown in FIG. 14, a construction information input window appears to analyze design change items within a new construction project, and the construction name, construction period, construction amount, type of construction, and other (select) can be entered through the input module. Meanwhile, when the search button is clicked, similar construction projects among the design change items of Daegu Urban Development Corporation can be searched to explore the design change details of the corresponding construction project.
[0067] When construction information is entered through the input module, as shown in Fig. 15, a design change history inquiry can be displayed with "Construction name: Noise reduction construction near Busan Expressway, Construction period: 1 year, Construction cost: 4 billion won, Construction type: Civil engineering work, Other (select)".
[0068] In addition, as shown in Fig. 16, design change details that occurred in construction projects of the Daegu Urban Development Corporation similar to the input construction information may also be provided. That is, as an output of such analysis results, "similar construction overview, design change items, reasons for design change, change costs, and reference document information" may be displayed, and when clicking on the output of design change document information used to generate the answer, the original document image as shown in Fig. 17 may also be viewed.
[0069] Next, the operation of a design change chatbot by a question-and-answer module according to the present invention will be described.
[0070] When a user clicks the Design Change Detection > Design Change Chatbot button in the state shown in FIG. 18, as shown in FIG. 19, the chatbot is automatically activated when moving to the chatbot screen with the user account, allowing the user to freely input questions related to design change documents and generate an answer by clicking the send button.
[0071] For example, as illustrated in FIG. 20, if a question is entered in the question input as "Tell me about the paving work performed by Daegu Urban Development Corporation," the question response module (600) can generate a response to the question and output the result.
[0072] Also, as shown in FIG. 21, clicking "View Reference Document" at the bottom of the response outputs information on the design change document used to generate the response, and clicking this reference document allows you to view the original document image as shown in FIG. 22.
[0073] As described above, the system and method for detecting risk factors in ultra-large AI-based technical documents according to the present invention can automatically detect risk factors, which are parts where design changes are expected, through automated document processing and information retrieval functions, thereby significantly improving work efficiency.
[0074] Although the invention made by the inventors has been specifically described according to the above embodiments, the invention is not limited to the above embodiments and can be modified in various ways without departing from the gist thereof.
[0075] By using the ultra-large AI-based technical document risk factor detection system and method according to the present invention, work efficiency can be significantly improved through automated document processing and information retrieval functions.
Claims
1. As a system that detects risk factors in technical documents based on mega-AI, A data collection module that collects analog and digital documents held within the organization, A text extraction module that automatically extracts text from analog and digital documents collected by the above data collection module through OCR and / or Text Block recognition, A module for constructing a data preprocessing pipeline for augmented search by dividing the text extracted from the above text extraction module into units recognizable by generative AI and vectorizing it, A design change detection module that automatically detects parts of a design document where design changes are expected based on design change history data collected from the above data collection module, A technical document risk factor detection system characterized by including a management server that transmits, receives, and manages information in conjunction with the above-mentioned data collection module, text extraction module, data preprocessing pipeline construction module, and design change detection module.
2. In Paragraph 1, A technical document risk factor detection system characterized by further including a storage module that stores meta-information extracted from the text extraction module and vector information from the preprocessing pipeline construction module.
3. In Paragraph 1, A technical document risk factor detection system characterized by the above-mentioned design change detection module detecting design change documents according to the project name, construction period, construction amount, type of construction, and options.
4. In Paragraph 3, The above design change detection module detects and outputs a history inquiry of similar construction design changes similar to the construction information entered by the user, and The above-mentioned similar construction design change history information includes information on the construction overview, design change items, reasons for design changes, change costs, and reference document information, and The above construction overview includes the project name, construction period, construction amount, type of construction, and contractor, and A technical document risk factor detection system characterized by the above reference document information being the original image of a design change.
5. In Paragraph 1, It further includes a question-response module that answers a question entered by a user using information about design changes detected by the above-mentioned design change detection module, and The above question-and-answer module provides reference documents for the response, and A technical document risk factor detection system characterized in that the above reference document is an original document image of design change document information used to generate a response.
6. As a method for detecting risk factors in technical documents based on mega-AI, (a) A data collection step for collecting analog and digital documents held within the organization, (b) A text extraction step for automatically extracting text from analog and digital documents collected in step (a) through OCR and / or Text Block recognition, (c) A step of constructing a data preprocessing pipeline for augmented search by dividing the text extracted in step (c) into units recognizable by generative AI and vectorizing it, (d) A design change detection step that automatically detects parts of a design document where design changes are expected based on design change history data collected in step (a) above, (e) A technical document risk factor detection system characterized by including a step of responding to a question entered by a user with information about design changes detected in step (d) above.
7. In Paragraph 6, The response in step (e) above provides a reference document, and A method for detecting technical document risk factors, characterized in that the above reference document is an original document image of design change document information used to generate a response.