Apparatus and method for providing corporate information based on ai-driven sentiment analysys

The information provision device uses a large-scale language model for sentiment analysis and anomaly detection to address the limitations of existing corporate monitoring systems, enabling real-time, quantitative evaluation and proactive risk management.

KR102997187B1Active Publication Date: 2026-07-29ANTOCK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
ANTOCK CO LTD
Filing Date
2025-09-26
Publication Date
2026-07-29

AI Technical Summary

Technical Problem

Corporate monitoring systems struggle to effectively analyze unstructured data from news and social media for real-time sentiment analysis and event detection, lacking sophisticated models for quantitative evaluation of business conditions and event detection.

Method used

An information provision device using a large-scale language model for sentiment analysis, which includes preprocessing, sentiment scoring, and anomaly detection to identify key documents and generate summary information for corporate monitoring.

Benefits of technology

Enables quantitative evaluation of business situations and rapid identification of critical issues through sentiment-based anomaly detection, providing meaningful summaries for proactive risk management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 112025110209611-PAT00050_ABST
    Figure 112025110209611-PAT00050_ABST
Patent Text Reader

Abstract

An information providing device is disclosed. An information providing device through large-scale language model-based sentiment analysis according to one embodiment comprises at least one processor; and at least one memory operatively connected to the at least one processor; wherein, upon execution, the at least one memory may store instructions that cause the at least one processor to receive a list of managed business operators, acquire a corpus regarding the target company through identification information of the target company included in the list of managed business operators, perform preprocessing on at least one document included in the corpus, extract at least one valid document containing a target keyword regarding the business operation of the target company from the at least one document, apply the at least one valid document to a sentiment analysis model to obtain a sentiment score for each of the at least one valid document, determine the priority of the at least one valid document based on the sentiment score of each of the at least one valid document, apply the at least one valid document to a large language model (LLM) to generate at least one LLM output, and provide summary information of the target company based on the priority.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The following embodiments relate to an apparatus and method for providing corporate information through sentiment analysis based on artificial intelligence, and more specifically, to an apparatus and method for providing information through sentiment analysis based on a large-scale language model. Background Technology

[0002] Corporate monitoring systems have primarily focused on analyzing structured data such as financial statements, public disclosures, and credit ratings, and have limitations in detecting changes in actual business conditions or abnormal signs in real time. While some systems have attempted to utilize unstructured data such as news and social media, most have been limited to simple keyword matching or rule-based filtering, making it difficult to detect meaningful events or perform quantitative analysis.

[0003] For example, there is a lack of sophisticated models capable of effectively processing the vast amount of unstructured text generated from news articles, and the absence of clear quantitative indicators for sentiment analysis or event detection makes it unsuitable for use as a real-time monitoring and alerting system for business status.

[0004] To address these issues, there is a growing need for technology capable of performing sentiment analysis, event detection, and time-series anomaly detection on article content, and inferring the state of a company based on a comprehensive quantitative model. The problem to be solved

[0005] The embodiments perform sentiment analysis and meaning-based summarization based on unstructured data corresponding to corporate-related content data, thereby quantitatively evaluating the management situation, risk indicators, positive and negative events of a specific company, and identifying the concentration of negative articles or the amount of rapid change at a specific point in time through sentiment score-based anomaly detection (e.g., Z-score analysis).

[0006] The embodiments can generate key points and sentiment-based keywords for each document through summary output using a large-scale language model, thereby enabling users or system administrators to quickly identify issues requiring priority review among a vast list of articles.

[0007] However, technical challenges are not limited to the technical challenges described above, and other technical challenges may exist. means of solving the problem

[0008] An information provision device based on large-scale language model sentiment analysis according to one embodiment comprises: at least one processor; and at least one memory operatively connected to the at least one processor; wherein, upon execution, the at least one memory may store instructions that cause the at least one processor to receive a list of managed business operators, acquire a corpus regarding said target company through identification information of said target company included in the list of managed business operators, perform preprocessing on at least one document included in the corpus, extract at least one valid document containing a target keyword regarding the business operation of said target company from said at least one document, apply said at least one valid document to a sentiment analysis model to acquire a sentiment score for each of said at least one valid document, determine the priority of said at least one valid document based on the sentiment score of each of said at least one valid document, apply said at least one valid document to a large language model (LLM) to generate at least one LLM output, and provide summary information of said target company based on said priority.

[0009] In one embodiment, the at least one memory may store instructions that, at execution, cause the at least one processor to determine the identification information including at least one of the target company’s Korean name, English name, English abbreviation, stock code, standard code, alias, representative name, or any combination thereof, generate a search keyword based on at least one of the Korean name, English name, English abbreviation, stock code, standard code, abbreviation, representative name, or any combination thereof, and apply the search keyword to an information providing platform regarding a news database, a digital content repository, or an online posting system to acquire the corpus.

[0010] In one embodiment, the at least one memory may store instructions that, at execution, cause the at least one processor to perform preprocessing on the at least one document included in the corpus by extracting a first comparison document and a second comparison document from the at least one document included in the corpus, applying the first comparison document and the second comparison document to a sentence transformer (SBERT) model trained to calculate semantic similarity between sentences, obtaining a document score regarding the content, structure, and form similarity of the first comparison document and the second comparison document, and removing one of the first comparison document or the second comparison document from the corpus based on the document score exceeding a predetermined threshold score.

[0011] In one embodiment, the at least one memory may store instructions that, at execution, cause the at least one processor to determine a first keyword including at least one of the Korean name of the target company, the English name of the target company, or any combination thereof, determine a second keyword for extracting a document created from a predetermined information providing platform from the at least one document, apply a third keyword predetermined as unrelated to the operation of the target company to the at least one document, remove a document containing the third keyword from the at least one document, and apply the first keyword and the second keyword to the at least one document from which the document containing the third keyword has been removed, thereby causing the at least one valid document to be extracted.

[0012] In one embodiment, the at least one memory may, at execution, allow the at least one processor to apply the at least one valid document to the sentiment analysis model to obtain a sentiment score for each of the at least one valid document, the sentiment score is normalized to a value included in a predetermined real number range, and includes criteria for classifying each of the at least one valid document as a negative document or a positive document, and the sentiment dictionary may include at least one sentiment keyword and a weight corresponding to the sentiment keyword. In one embodiment, the at least one memory may, at execution, allow the at least one processor to preprocess each of the at least one valid document to extract feature vectors, apply the feature vectors to the sentiment analysis model which is composed of a structure in which a plurality of layers are stacked and includes Multi-head Attention and Position-wise Feed Forward Network in the plurality of layers, obtain a sentiment score based on an attention operation, and classify the sentiment score to estimate at least one of the positive document or negative document for each of the at least one valid document. there is.

[0013] In one embodiment, the at least one memory may store instructions that, at execution, cause the at least one processor to determine a first average sentiment score based on the sentiment score of each of the at least one valid document, determine a second average sentiment score based on the sentiment score determined at a preceding time prior to the time of acquiring the sentiment score of each of the at least one valid document, and determine the priority of the at least one valid document based on the first average sentiment score, the second average sentiment score, and the standard deviation determined based on the sentiment score determined at the preceding time.

[0014] In one embodiment, the at least one memory may store instructions that, at execution, cause the at least one processor to determine the Z-score of the at least one valid document based on the first average sentiment score, the second average sentiment score, and the standard deviation, and if the Z-score is not included in a predetermined score range, determine the priority of the at least one valid document as a first priority for priority review, and if the Z-score is included in a predetermined score range, determine the priority of the at least one valid document as a second priority for general review.

[0015] In one embodiment, the LLM output may include at least one of the summary keywords included in the at least one valid document and the sentiment score according to a predetermined prompt.

[0016] In one embodiment, the at least one memory may store instructions that cause the at least one processor to identify the past identification information, based on a change history between the past identification information corresponding to the target company and the past identifier of the target company and the current identification information corresponding to the current identifier of the target company, so as to include the at least one document corresponding to the past identification information in the corpus.

[0017] A method for providing information through sentiment analysis based on a large-scale language model according to one embodiment may include: an operation of obtaining a corpus regarding a target company through identification information of the target company included in the list of management business operators based on receiving a list of management business operators; an operation of extracting at least one valid document containing a target keyword regarding the business operation of the target company from at least one document based on performing preprocessing on at least one document included in the corpus; an operation of applying the at least one valid document to a sentiment analysis model to obtain a sentiment score for each of the at least one valid document; an operation of determining the priority of the at least one valid document based on the sentiment score of each of the at least one valid document; and an operation of providing summary information of the target company based on at least one LLM output generated by applying the at least one valid document to a large language model (LLM) and the priority. Effects of the invention

[0018] The embodiments quantitatively evaluate a specific company's business situation, risk indicators, positive and negative events, etc., by performing sentiment analysis and meaning-based summarization based on documents corresponding to corporate-related content, and can identify the concentration of negative articles or the amount of rapid change at a specific point in time through sentiment score-based anomaly detection (e.g., Z-score analysis).

[0019] The embodiments can generate key points and sentiment-based keywords for each article through summary output using a large-scale language model, thereby enabling users or system administrators to quickly identify issues requiring priority review among a vast list of articles. Brief explanation of the drawing

[0020] FIG. 1 shows a schematic block diagram of an information providing device according to one embodiment. FIG. 2 shows a flowchart of an information provision method according to one embodiment. FIG. 3 shows a flowchart of a method for providing summary information about a target company from a corpus in an information providing device according to one embodiment. FIG. 4 illustrates an example of a method for obtaining a corpus from an information provision platform based on identification information of a target company in an information provision device according to one embodiment. FIG. 5 illustrates an example of a preprocessing method for removing similar documents based on applying documents included in a corpus to a sentence transformer model in an information providing device according to one embodiment. FIG. 6 shows a flowchart of a method for extracting at least one valid document from a corpus that has undergone preprocessing in an information providing device according to one embodiment. FIG. 7 illustrates an example of a method for obtaining a Z-score based on applying at least one valid document to a sentiment analysis model in an information providing device according to one embodiment. FIG. 8 shows an example of a method for expressing an emotional score in an information providing device according to one embodiment. FIG. 9 shows an example of visualizing a change in sentiment score based on the time of a target company's application for rehabilitation in an information providing device according to one embodiment. FIG. 10 shows an example of a user interface that outputs summary information provided based on LLM output and priority in an information providing device according to one embodiment. FIG. 11 is a drawing illustrating an example of an emotional dictionary implementation according to one embodiment. Specific details for implementing the invention

[0021] Specific structural or functional descriptions of embodiments according to the concept of the present invention disclosed herein are provided merely for the purpose of explaining embodiments according to the concept of the present invention, and embodiments according to the concept of the present invention may be implemented in various forms and are not limited to the embodiments described herein.

[0022] Embodiments according to the concept of the present invention may be subject to various modifications and may take various forms; therefore, embodiments are illustrated in the drawings and described in detail in this specification. However, this is not intended to limit the embodiments according to the concept of the present invention to specific disclosed forms, and includes modifications, equivalents, or substitutions that fall within the spirit and scope of the present invention.

[0023] Terms such as "first" or "second" may be used to describe various components, but said components should not be limited by said terms. For the sole purpose of distinguishing one component from another, for example, without departing from the scope of rights according to the concept of the present invention, the first component may be named the second component, and similarly, the second component may be named the first component.

[0024] When it is stated that one component is "connected" or "connected" to another component, it should be understood that while it may be directly connected or connected to that other component, there may also be other components in between. Conversely, when it is stated that one component is "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between. Expressions describing the relationships between components, such as "between," "exactly between," or "directly adjacent to," should be interpreted in the same way.

[0025] The terms used herein are used merely to describe specific embodiments and are not intended to limit the invention. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this specification, terms such as “comprising” or “having” are intended to specify the existence of the described features, numbers, steps, actions, components, parts, or combinations thereof, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0026] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as generally understood by those skilled in the art to which the present invention pertains. Terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and should not be interpreted in an ideal or overly formal sense unless explicitly defined in this specification.

[0028] In this specification, a module may mean hardware capable of performing functions and operations according to each name described in this specification, computer program code capable of performing specific functions and operations, or an electronic recording medium loaded with computer program code capable of performing specific functions and operations, such as a processor or a microprocessor.

[0029] In other words, a module may mean a functional and / or structural combination of hardware for carrying out the technical concept of the present invention and / or software for driving said hardware.

[0031] Hereinafter, embodiments will be described in detail with reference to the attached drawings. However, the scope of the patent application is not limited or restricted by these embodiments. Identical reference numerals in each drawing indicate identical components.

[0033] FIG. 1 shows a schematic block diagram of an information providing device according to one embodiment.

[0034] The information providing device (100) described in this document is a system comprising hardware and software components for providing summary information of a target company through large-scale language model-based sentiment analysis, and may include a processor (110), memory (120), and a network interface (130). The information providing device (100) may be implemented as various types of computing devices, such as, for example, a general GPU or TPU-based learning server, a cloud-based machine learning platform, or an on-device learning environment.

[0035] Specifically, the information providing device (100) can automatically collect corporate-related content data and perform corporate management status evaluation and risk management through sophisticated data processing and analysis based thereon. In particular, the information providing device (100) can collect daily content data through an automated module based on Robotic Process Automation (RPA), and by processing the collected unstructured data according to a multi-stage analysis pipeline including preprocessing, structuring, sentiment analysis, and identification of valid articles, it can detect credit risk, management stability, or signs of issues of the target company at an early stage. Through this, the information providing device (100) can go beyond simply providing articles to implement a real-time monitoring system based on data science and provide meaningful information for corporate analysis and proactive risk prevention. For example, a person skilled in the art will understand that the content data may include arbitrary information related to the target company, such as news articles, corporate announcements, and information related to public institutions.

[0036] The information providing device (100) may be implemented as a printed circuit board (PCB), such as a motherboard, an integrated circuit (IC), or a system on chip (SoC). For example, the information providing device (100) may be implemented as an application processor.

[0037] Additionally, the information providing device (100) may be implemented in a PC (personal computer), a data server, or a portable device.

[0038] Portable devices can be implemented as laptop computers, mobile phones, smartphones, tablet PCs, mobile internet devices (MID), personal digital assistants (PDA), enterprise digital assistants (EDA), digital still cameras, digital video cameras, portable multimedia players (PMP), personal navigation devices (PND), handheld game consoles, e-books, or smart devices. Smart devices can be implemented as smart watches, smart bands, or smart rings.

[0039] The processor (110) can process data stored in memory (120). The processor (110) can execute computer-readable code (e.g., software) stored in memory (120) and instructions triggered by the processor (110).

[0040] The "processor (110)" may be a data processing device implemented in hardware having a circuit having a physical structure for executing desired operations. For example, the desired operations may include code or instructions included in a program.

[0041] For example, a data processing device implemented in hardware may include a microprocessor, a central processing unit, a processor core, a multi-core processor, a multiprocessor, an Application-Specific Integrated Circuit (ASIC), and a Field Programmable Gate Array (FPGA).

[0042] The memory (120) can store instructions (or programs) executable by the processor (110). For example, the instructions may include instructions for executing the operation of the processor and / or the operation of each component of the processor.

[0043] The memory (120) can be implemented as a volatile memory device or a non-volatile memory device.

[0044] Volatile memory devices can be implemented as DRAM (dynamic random access memory), SRAM (static random access memory), T-RAM (thyristor RAM), Z-RAM (zero capacitor RAM), or TTRAM (Twin Transistor RAM).

[0045] Non-volatile memory devices can be implemented as EEPROM (Electrically Erasable Programmable Read-Only Memory), flash memory, MRAM (Magnetic RAM), Spin-Transfer Torque (STT)-MRAM, Conductive Bridging RAM (CBRAM), FeRAM (Ferroelectric RAM), PRAM (Phase change RAM), Resistive RAM (RRAM), Nanotube RRAM, Polymer RAM (PoRAM), Nano Floating Gate Memory (NFGM), holographic memory, Molecular Electronic Memory Device, or Insulator Resistance Change Memory.

[0047] FIG. 2 shows a flowchart of an information provision method according to one embodiment.

[0048] A processor according to one embodiment (e.g., processor (110) of FIG. 1) can obtain a corpus regarding a target company through identification information of a target company included in a list of management companies based on receiving a list of management companies in operation S210.

[0049] For example, a list of managed operators may refer to a list containing information regarding multiple companies (e.g., operators) subject to surveillance or monitoring according to predefined criteria. Specifically, the list of managed operators may be designated by a managing entity, such as a financial institution, credit card company, investment firm, insurance company, or user, or may be generated periodically based on automated filtering logic, and may be utilized as basic information to determine the scope of the group of companies subject to monitoring.

[0050] For example, a target company may refer to one or more specific companies included in the list of managed operators that are the direct targets of monitoring. Specifically, the target company functions as the entity for which sentiment-based content data analysis is performed from various perspectives, such as credit risk, changes in liquidity, financial stability, and business events; generally, this may include listed companies, unlisted small and medium-sized enterprises, and partner companies.

[0051] For example, identification information refers to information used to clearly identify a target company in external databases or news search systems, and may include, for instance, the Korean company name, English company name, English abbreviation, stock code, standard code, unique corporate identifier, representative name, registration number, etc. Specifically, one or more pieces of identification information may be used in combination to enhance the identification accuracy of the target company.

[0052] Specifically, the processor can identify past identification information based on the change history between past identification information regarding the target company’s past identifier and current identification information regarding the target company’s current identifier, corresponding to the target company, so as to include at least one document corresponding to the past identification information in the corpus. That is, the identification information may include past identification information and current identification information.

[0053] For example, historical identification information refers to identifiers previously used by the target company and may be information that is no longer in use due to reasons such as changes in the target company's name, trade name, mergers and acquisitions (M&A), or changes in the representative. Historical identification information may include past Korean trade names, English trade names, English abbreviations, stock codes, standard codes, acronyms, and representative names. Historical identification information may still be in use in publications, news articles, or disclosure materials written in the past. Therefore, the present invention enables the construction of a corpus without information omission, despite changes in the time series of the target company, by extracting and identifying historical identification information.

[0054] For example, current identification information refers to the identifier currently used by the target company at the present time, and may refer to the latest name or code that serves as a standard for identifying the target company. Current identification information may include information such as the currently used Korean company name, English company name, English abbreviation, stock code, standard code, abbreviation, or representative name. In this invention, the target company is identified based on current identification information, and by referencing change history information to include documents based on past identification information, the corpus is expanded to enable time-series corporate analysis without omissions. For example, the processor can collect valid company name information by checking whether the company name of the target company identified based on current identification information is closed, and if a company name is closed, including each closed company name as past identification information. In addition, the processor can collect additional identification information regarding the target company by searching for current / past company names collected within a site capable of collecting public data in conjunction with the representative name.

[0055] According to an additional embodiment, if the company name corresponding to the database-stored identification information of the target company contains English letters, the processor may additionally store the result converted into a Korean business name by mapping a fixed Korean phonetic pronunciation corresponding to the English letters, or conversely, if the company name contains a Korean phonetic transcription corresponding to the English letters, the processor may additionally store the result converted into a Korean business name by mapping a fixed English letter corresponding to the Korean phonetic transcription. For example, if the previously stored company name is "LA Fire Insurance," the processor may additionally database "LA Fire Insurance," which maps the Korean phonetic transcription corresponding to the English letters, as the identification information of the target company; conversely, if "LA Fire Insurance" is already stored as identification information, the processor may diversify and database the identification information for the target company by converting and storing "LA Fire Insurance" as additional identification information corresponding to the target company.

[0056] According to an additional embodiment, the processor can diversify identification information for a target company through the expansion of abbreviations corresponding to company names. For example, if an official company name includes an abbreviation, the processor can database the full name of the abbreviation by mutually mapping it to the official company name. The processor can database the identification information of the target company not only for the official company name but also for each name searched on various portals corresponding to the official company name, by mapping the full name of the abbreviation included in the name to the searched name. For example, if an official company name "BGI Busan Guarantee" exists, and BGI included in the name is an abbreviation for Busan Guarantee Insurance, the processor can database the identification information for the same company by mapping "Busan Guarantee Insurance" to "BGI Busan Guarantee."

[0057] In an additional embodiment, if the processor determines that the expression "Hanshin" is duplicated among distinct companies such as "Hanshin Financial Group," "Hanshin Bank," and "Hanshin Networks," but each company corresponds to a different company, the processor may not determine "Hanshin" as corresponding to the abbreviation or representative name of the company, and may database the company name corresponding to the separate company without mutually mapping the individual companies.

[0058] For example, a corpus may refer to a dataset containing a large amount of unstructured document data collected from any accessible database, such as news search platforms, digital content repositories, online publishing systems, or public institution websites, utilizing identification information. The corpus may include various articles, press releases, blog posts, or disclosure documents related to the target company, and may include source data for analysis used in subsequent processing such as sentiment analysis, keyword extraction, and LLM summarization.

[0059] Specifically, the corpus may include at least one of article-like data in the form of unstructured text, disclosures, reports, presentation materials, or any combination thereof; however, it will be understood by a person skilled in the art that the types of data included in the corpus are not limited to the examples presented and may include any data collected through various platforms.

[0060] Specifically, the processor can prevent unnecessary data collection and efficiently allocate analysis resources by selectively targeting only the group of companies of interest to the managing entity, and by performing the process of receiving the list of managed operators and collecting the corpus periodically or on an event-based basis, it can rapidly detect changes in management or the occurrence of risks in target companies.

[0061] A detailed explanation of the operation to acquire a corpus regarding the target company is described later in Figure 4 below.

[0062] Based on performing preprocessing on at least one document included in the corpus in operation S220, the processor can extract at least one valid document containing a target keyword regarding the business operation of the target company from at least one document.

[0063] For example, at least one document may refer to text data or digital content collected containing information about the target company. At least one document may include various forms of unstructured text, such as news articles, press releases, public disclosures, online posts, and blogs, and the degree of structuring may vary depending on the time of collection, the source of creation, and the format. At least one document is processed as an analysis unit within the entire corpus and may be utilized for sentiment analysis and summarization.

[0064] For reference, for convenience of explanation, each document included in at least one document in this specification may be described as being identical to at least one of corporate news, daily news information, news articles, news documents, or any combination thereof. The types presented above as examples of documents are listed for illustrative purposes only, and it will be understood by those skilled in the art that the documents subject to the present invention are not limited to the presented examples and any document / content may be subject to the invention.

[0065] For example, preprocessing may refer to text cleaning and processing performed to derive meaningful analysis results from collected documents. Specifically, preprocessing may include treatments such as stopword removal, morphological analysis, named entity recognition (NER), sentence segmentation, duplicate document removal, or context-based normalization.

[0066] A detailed description of the operation to perform preprocessing on at least one document is provided below in Fig. 5.

[0067] For example, target keywords may include key expressions or concepts within documents that are identified as indicators of the target company's business activities or risks. For instance, target keywords may include negative or positive events directly related to business operations, such as filing for rehabilitation, restructuring, poor performance, financial difficulties, or business suspension. Target keywords can be derived from a predefined keyword dictionary or a machine learning-based keyword extraction model.

[0068] For example, at least one valid document may refer to a preprocessed document that contains the target keyword and whose content is determined to have substantial relevance to the business activities of the target company. At least one valid document is utilized as input data for NLP processing such as sentiment analysis, scoring, and summarization.

[0069] Specifically, the processor can ensure the quality of data delivered to the sentiment analysis and summarization stages and provide the effect of reducing noise caused by an excessive amount of information by removing promotional articles, duplicate articles, and simple reference articles that are not directly related to the business activities of the target company.

[0070] A detailed description of the operation to extract at least one valid document containing a target keyword from at least one document is provided below in Fig. 6.

[0071] In operation S230, the processor can obtain a sentiment score for each of at least one valid document by applying at least one valid document to a sentiment analysis model that performs sentiment classification by deriving labels for related sentiments. More specifically, the processor can calculate a final sentiment score for the results classified through the sentiment analysis model by assigning weights based on a sentiment dictionary.

[0072] For example, a sentiment dictionary can refer to a database or reference table that defines the emotional tendencies of words or phrases contained within text data. Each word or phrase is generally classified into one of three sentiment values—positive, negative, or neutral—and additionally, a weight may be assigned based on the intensity of the sentiment. For instance, keywords such as "gratitude," "respect," and "joy" can be defined in the dictionary as positive sentiment keywords, while keywords such as "shock," "regret," and "disgust" can be defined as negative sentiment keywords. Sentiment dictionaries can be configured as domain-specific versions to improve analysis accuracy and can be utilized as a standard for judging sentiment within text.

[0073] For example, a sentiment analysis model refers to an algorithm or learning-based model that takes text data as input and automatically classifies or quantifies emotional meanings. The sentiment analysis model in the present invention can determine the sentiment tendency of an entire document by comprehensively considering the occurrence and frequency of specific words and contextual usage patterns based on a sentiment dictionary. The sentiment analysis model may be a rule-based configuration or a deep learning-based model composed of multiple attention layers and a feedforward network. The sentiment analysis model can be utilized for operations such as classifying each valid document as positive, negative, or neutral, or generating a quantitative sentiment indicator (i.e., a sentiment score). For example, the sentiment analysis model is implemented as a neural network-based structure including multi-head attention and a feedforward network, and the sentiment dictionary can be utilized as auxiliary information used to determine the weights of key input words. A detailed description of the sentiment analysis model and the sentiment dictionary is provided below in Fig. 7.

[0074] For example, a sentiment score refers to a quantitative sentiment indicator produced as a result of processing by a sentiment analysis model, and can be expressed by quantifying the direction (positive or negative) and strength of the sentiment of each valid document. Sentiment scores are generally provided as normalized values ​​within a real number range, and, for example, can be set to a range of -2 to +2 or 0 to 1. Specifically, a higher sentiment score indicates a stronger positive meaning, while a lower score indicates a stronger negative sentiment.

[0075] Specifically, the processor can express whether each valid document has a positive or negative impact on the target company's business activities using quantified indicators, and can perform rapid and objective priority determination, risk assessment, and alert trigger determination.

[0076] A detailed explanation of the action for obtaining the emotional score is described later in Fig. 7 below.

[0077] The processor can determine the priority of at least one valid document based on the sentiment score of each of at least one valid document in operation S240.

[0078] For example, the processor may classify a document as a 'priority document' if the sentiment score deviates by more than a certain threshold from a predefined average value, and classify it as a general review otherwise. The processor may utilize the mean and standard deviation of the sentiment scores calculated based on the distribution of sentiment scores by company over a specific period (e.g., the last 30 or 180 days).

[0079] For example, the processor may assign a review grade to valid documents by subdividing them into a High Risk Group if the Z-score is ±2.0 or higher, a Watched Group if it is ±1.5 or higher but less than 2.0, and a General Group otherwise. Additionally, if there is a sharp drop in the sentiment score or if negative sentiment is concentrated in multiple articles at the same time, the processor may trigger an Alert status for the target company.

[0080] The processor can provide summary information of the target company based on at least one LLM output and priority generated by applying at least one valid document to a large language model (LLM) in operation S250.

[0081] For example, a large language model may refer to a natural language processing-based artificial intelligence model capable of learning a large natural language corpus (i.e., a corpus) to understand the linguistic meaning of input text and generate appropriate text output. The LLM is a deep learning-based structure containing parameters and can be designed based on a transformer-family model architecture. In the present invention, the large language model can receive the content of a valid document as input, summarize the context or extract key information, and return the business status of the company, whether an event has occurred, sentiment trends, etc., in a summarized form.

[0082] For example, LLM output may refer to a text response generated as a result of inputting valid documents into a large language model. LLM output may include summary keywords and sentiment scores contained in at least one valid document according to a predetermined prompt. Specifically, LLM output may be generated in a structured form according to a predefined prompt and may include, for example, natural language sentences or token sequences in a form corresponding to specific requests such as "article summary," "key keywords," "sentiment judgment result," or "risk indication summary."

[0083] For example, summary information may refer to a unit of information that concisely expresses the target company's management status, risk events, recent issues, etc., generated based on LLM output related to valid documents. Summary information is provided to help general users or monitoring personnel quickly understand the company's status and may consist of sentence-based summaries or keyword lists, for instance, such as "Surge in negative articles regarding recent poor performance," "Confirmation of bankruptcy filing," or "Ongoing reports on restructuring."

[0084] A detailed explanation regarding the summary information of the target company is described later in Figure 10 below.

[0086] FIG. 3 shows a flowchart of a method for providing summary information about a target company from a corpus in an information providing device according to one embodiment.

[0087] A processor according to one embodiment can collect corporate news in operation S310. Specifically, the processor can collect a corpus containing article-like data as corporate news by applying search keywords determined based on identification information of a target company to an information providing platform.

[0088] For example, the processor may generate search keywords through a combination of identification information from the list of companies to be queried ("company name," "company name and representative name," or "stock code," etc.). In the case of a company name, the processor may perform querying and collection by generating derivatives such as an English company name, a company name combined in English and Korean, or a stock code. Here, the scope of the collection period may include daily news information prior to a predetermined date immediately preceding the date of inquiry.

[0089] The processor can remove similar articles included in the collected corporate news in operation S320. Specifically, the processor can remove similar articles based on the document score obtained by applying the collected corporate news to a sentence transformer model.

[0090] For example, a sentence transformer model is a pre-trained deep learning language model designed to quantitatively calculate semantic similarity between sentences, and can be utilized to efficiently generate sentence embedding vectors, particularly in the field of natural language processing (NLP). A sentence transformer model is based on pre-trained transformer structures such as BERT (Bidirectional Encoder Representations from Transformers) or RoBERTa, and can be configured to generate high-dimensional embedding vectors for input sentences or documents and to quantify semantic similarity between sentences or documents by comparing cosine similarity between embedding vectors.

[0091] For example, since the Sentence Transformer model is trained so that sentences with identical or similar meanings have high similarity scores, it can provide various applications such as determining article duplication, filtering semantic overlap, and semantic clustering. The Sentence Transformer model receives collected article data as input and calculates a document score representing the semantic similarity between each document; if the document score exceeds a predetermined threshold, the article is considered semantically duplicated from an existing article and can be classified as a target for removal.

[0092] The processor can train a sentence transformer model. For example, the sentence transformer model may include a neural network. The neural network may include multiple layers, and each layer may include multiple nodes. A node may have a node value determined based on an activation function. A node in any layer may be connected to a node in another layer (e.g., another node) through a link (e.g., a connection edge) having a connection weight. A node's node value may be propagated to other nodes through the link. In the inference operation of the neural network, node values ​​may be forward propagated from the previous layer to the next layer.

[0093] For example, in a sentence transformer model, the forward propagation operation can represent an operation that propagates node values ​​based on input data from the input layer toward the output layer of the sentence transformer model. That is, the node value of a given node can be propagated (e.g., forward propagation) to the node in the next layer (e.g., the next node) connected to the node via a connection line. For instance, a node can receive a value weighted by connection weights from previous nodes (e.g., multiple nodes) connected via a connection line.

[0094] The node value of a node can be determined based on applying an activation function to the sum of weighted values ​​received from previous nodes (e.g., weighted sum). The parameters of the neural network may include, for example, the aforementioned connection weights. The parameters of the neural network can be updated so that the objective function value described below changes in the targeted direction (e.g., the direction in which loss is minimized).

[0095] A learned sentence transformer model may represent a model learned through machine learning and may be a learned machine learning model that outputs a training output from a training input. A machine learning model (e.g., a learned sentence transformer model) may be generated through machine learning. The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above.

[0096] In the case of supervised learning, the machine learning model described above can be trained based on training data containing pairs of training inputs and training outputs mapped to those training inputs. For example, the machine learning model can be trained to output a training output from a training input. During training, the machine learning model can generate a temporary output in response to the training input and can be trained to minimize the loss between the temporary output and the training output (e.g., the target of training). During the training process, the parameters of the machine learning model (e.g., connection weights between nodes / layers in a neural network) can be updated according to the loss. This training can be performed, for example, on the information providing device where the machine learning model is performed (e.g., the information providing device (100) of FIG. 1) itself, or through a separate server. The machine learning model that has completed training (e.g., the trained sentence transformer model) can be stored in memory (e.g., the memory (120) of FIG. 1).

[0097] In operation S330, the processor can determine valid articles from corporate news from which similar articles have been removed. Specifically, the processor can determine that news articles containing the target keyword of the target company are valid articles from the corporate news from which similar articles have been removed.

[0098] For example, the processor can determine valid articles by determining whether articles collected based on search keywords are actually related to a company or its management. Specifically, the processor can remove articles published for purposes such as promotion or advertising, geographical names (e.g., landmarks), or simple references. Here, the processor can extract valid articles by removing from the corpus articles that include the target company's trade name and keywords related to obituaries, weddings, or personnel announcements.

[0099] The processor can perform sentiment analysis on valid articles in operation S340. Specifically, the processor can obtain sentiment results corresponding to the sentiment score based on the sentiment score obtained by applying the valid articles to a sentiment analysis model. Here, the sentiment results include positive, negative, or neutral, and can be stored in a form paired with the valid articles.

[0100] The processor can determine whether to prioritize review a valid article for which sentiment analysis has been performed in operation S350. Specifically, the processor can determine whether a valid article is subject to priority review based on statistical figures of the valid article for which sentiment analysis has been performed. For example, the processor can determine whether to prioritize review a valid article for which sentiment analysis was performed in operation S340 based on a comparison of the sentiment scores of valid articles of the target company performed in the past (e.g., mean or standard deviation) with the sentiment scores of the valid article for which sentiment analysis was performed in operation S340.

[0101] The processor can provide the user with an LLM summary for valid articles selected for priority review in operation S360. Specifically, the processor can provide the user with summary information of the target company based on the priority and LLM output generated by applying the valid articles to a large language model.

[0103] FIG. 4 illustrates an example of a method for obtaining a corpus from an information provision platform based on identification information of a target company in an information provision device according to one embodiment.

[0104] A processor according to one embodiment (e.g., processor (110) of FIG. 1) can determine identification information (410) including at least one of the target company's Korean name, English name, English abbreviation, item code, standard code, abbreviation, representative name, or any combination thereof.

[0105] The processor can generate search keywords (420) based on the identification information (410) determined. Specifically, the processor can determine identification information (410) for target companies identified through the list of managed business operators. The processor can generate search keywords (420) based on the determined identification information (410). The search keywords (420) may be text strings composed of a single or multiple combination of one or more identification information, such as a form combining a Korean company name ("Ganada") and an English abbreviation ("GND") ("Ganada GND"), or a form including a stock code ("3732255") and a representative name ("Kim Cheol-su"). Additionally, the search keywords (420) may include Boolean operators (AND, OR, etc.) or be generated in the form of wildcards or regular expressions, and are not limited to configurations intended to increase the data search efficiency of the information provision platform (430).

[0106] Search keywords (420) can be applied to information provision platforms (430), such as news databases, digital content repositories, or online publishing systems. A processor can acquire a corpus (440) containing article-like data (e.g., news, announcements, blogs, etc.) associated with the search keywords (420). Such keyword generation and corpus acquisition enable more accurate and comprehensive news collection by taking into account the diversity of corporate identification information and the unstructured nature of news sources.

[0108] FIG. 5 illustrates an example of a preprocessing method for removing similar documents based on applying documents included in a corpus to a sentence transformer model in an information providing device according to one embodiment.

[0109] A processor according to one embodiment (e.g., processor (110) of FIG. 1) can perform preprocessing to remove similar documents included in the corpus (510) based on applying at least one document included in the corpus (510) to a sentence transformer model (540).

[0110] Specifically, the processor can extract a first comparison document (520) and a second comparison document (530) from at least one document included in the corpus (510). However, the method by which the processor extracts comparison documents is not limited to this, and multiple comparison documents may be extracted.

[0111] The processor can apply the first comparison document (520) and the second comparison document (530) to a sentence transformer model (540) trained to calculate the semantic similarity between sentences, and obtain a document score (550) regarding the content, structure, and form similarity of the first comparison document (520) and the second comparison document (530).

[0112] For example, the document score (550) is a numerical value representing the semantic similarity between two or more documents, which can be calculated by reflecting the content, structure, and form elements of each document. In this specification, the document score (550) may refer to a value calculated through cosine similarity or other distance-based similarity functions (e.g., Euclidean distance, Manhattan distance, etc.) between document embedding vectors generated by the sentence transformer model.

[0113] The document score (550) can be expressed by the following mathematical formula 1:

[0114]

[0115] Here, can mean document score (550), and may mean the feature vector of the first comparison document (520), and can mean the feature vector of the second comparison document (530).

[0116] For example, the document score (550) is expressed as a continuous real value between -1 and 1, and a value closer to 1 indicates that the two documents are semantically very similar, a value closer to 0 indicates a low semantic association, and a value closer to -1 indicates that the meanings are opposite. The document score (550) can be used as a key criterion for determining duplicates between articles, classification based on content similarity, document clustering, etc., and in this specification, by removing one document for document pairs that exceed a certain threshold, it contributes to increasing the reliability of the analysis and the efficiency of data processing.

[0117] The processor may perform preprocessing on at least one document included in the corpus (510) by removing one of the first comparison document (520) or the second comparison document (530) from the corpus (510) based on the fact that the document score (550) exceeds a predetermined threshold score. According to an embodiment, the processor may perform preprocessing in the process of removing duplicate documents by retaining the most recent document and removing documents that were created earlier and whose similarity is greater than or equal to a predetermined threshold.

[0119] FIG. 6 shows a flowchart of a method for extracting at least one valid document from a corpus that has undergone preprocessing in an information providing device according to one embodiment.

[0120] A processor according to one embodiment (e.g., processor (110) of FIG. 1) can determine, in operation S610, a first keyword including at least one of the Korean name of a target company, the English name of a target company, or any combination thereof.

[0121] For example, the first keyword is a keyword for identifying a target company and may refer to an element used to improve mutual identification accuracy. The first keyword includes at least one of the target company's Korean name, the target company's English name, or any combination thereof, and may be used by a search system or data collection system to efficiently search for and classify articles related to the company. The first keyword may be used to identify valid documents among the documents included in the corpus that have a direct association with the target company.

[0122] The processor can determine a second keyword for extracting a document created from a predetermined information providing platform in at least one document in operation S620.

[0123] For example, the second keyword refers to a keyword used to distinguish specific sources among articles collected from an information provision platform for collecting article data, such as, for example, a news database, a digital content repository, or an online publishing system. The second keyword may include predefined media company names or platform identification information to selectively extract documents from highly reliable information sources. For example, the second keyword may include keywords containing media company names such as "Chosun Ilbo," "Yonhap News," or "KBS," and may be configured to minimize noise and increase the accuracy of information during the data collection process.

[0124] In operation S630, the processor can remove documents containing the third keyword from at least one document by applying a third keyword, which is predetermined to be unrelated to the operation of the target enterprise, to at least one document.

[0125] For example, the third keyword is a filtering criterion established to identify and remove documents unrelated to the company's actual business activities, and may include words or phrases used for social events, advertisements, personnel changes, family events, or simple mentions that are irrelevant to the target company. For instance, since keywords such as "obituary," "award," "marriage," "personnel appointment," and "appointment" are typically found in articles that do not reflect the actual operational status of the company, the processor can ensure the quality of valid documents and improve the accuracy of business status analysis by removing documents containing the third keyword related to the aforementioned keywords.

[0126] In operation S640, the processor can extract at least one valid document by applying the first keyword and the second keyword to at least one document from which a document containing the third keyword has been removed.

[0127] For example, the processor can extract valid documents from a set of documents removed including a third keyword by applying together a first keyword for identifying the target company (e.g., "Ganada", "GANADA") and a second keyword representing a predefined reliable information provider platform (e.g., "Chosun Ilbo", "Yonhap News"). For example, if "GANADA" and "Chosun Ilbo" exist simultaneously in the document title or body, the document may be considered related to the business activities of the target company and classified as a valid document.

[0128] The processor may apply the first keyword and the second keyword individually to assign a score to each document and determine whether the document is valid based on the weighted average of the scores or a combined value based on calculations. For example, if the first keyword match rate is 80% or higher and the source reliability of the second keyword exceeds a certain standard, the processor may consider the document to be valid.

[0129] The processor can improve the accuracy and reliability of sentiment analysis and LLM summaries by removing documents that may be considered unnecessary or noise through operations S610 to S640, and by selecting only reliable articles directly related to the activities of the target company. In particular, by applying a combination of the first, second, and third keywords rather than simply relying on keyword matching, the sophistication of article filtering is enhanced, and the effectiveness of early detection of business status and risk prediction can be increased.

[0130] Additionally, the processor may vectorize the meaning of the entire sentence of the target document based on a pre-trained natural language processing-based learning model (e.g., KeyBERT), generate a core keyword list by extracting keywords whose similarity to the entire meaning of the target document exceeds a predetermined threshold, and determine the final valid document by performing additional filtering based on the generated keyword list. More specifically, the processor may quantitatively evaluate the similarity between the keywords extracted through the pre-trained natural language processing-based learning model and the meaning of the target document, and generate a core keyword list based on keywords that exceed a preset threshold. The processor may determine the target document as a valid document if i) the target company's 'business name' is included in the generated core keyword list. ii) The processor may determine the target document as a valid document if a predefined keyword related to business activities (e.g., mergers and acquisitions, attracting investment, new business, strategic alliance, bankruptcy, filing for rehabilitation, etc., but a person skilled in the art will understand that the business-related keyword is not limited to the examples presented) is included in the core keyword list, and iii) if the 'company name' is not included in the core keyword list and only the 'representative name' is included, the processor may determine the target document as a valid document only if the 'company name' exists within the body of the target document or if the association with the 'company name' is above a pre-set threshold.

[0131] FIG. 7 illustrates an example of a method for obtaining a Z-score based on applying at least one valid document to a sentiment analysis model in an information providing device according to one embodiment.

[0132] A processor according to one embodiment (e.g., processor (110) of FIG. 1) can produce at least one sentiment score (740) by inputting at least one valid document (710) into a sentiment analysis model (720).

[0133] According to an embodiment, the processor (110) may perform preprocessing operations such as removing HTML tags, cleaning up whitespace, and merging sentences before the valid document (710) is input into the sentiment analysis model (720).

[0134] A processor (110) according to one embodiment can obtain an output including a label and a probability corresponding to the label by inputting at least one valid document (710) that has been preprocessed into a pre-trained LLM-based sentiment analysis model (720). The label included in the output may be information about the sentiment keyword to which the valid document (710) belongs (e.g., astonishment, sadness, surprise, respect, gratitude, welcome, etc.), and the probability may correspond to the probability that the valid document (710) corresponds to each label. For example, the output of the sentiment analysis model (720) may include a plurality of labels (emotions) corresponding to the valid document (710) and a probability corresponding to each emotion.

[0135] The sentiment analysis model (720) is configured with a structure in which multiple layers are stacked, and can be implemented as a large language model (LLM) that includes Multi-head Attention and Position-wise Feed Forward Network in the multiple layers. Additionally, the sentiment analysis model (720) may be a model that has been pre-trained to generate the previously described output based on an input corresponding to at least one of a valid document, a result of preprocessing a valid document, and a feature vector extracted from a valid document. For the sake of explanation, the sentiment analysis model (720) is described below as generating an output using the result of preprocessing a valid document (710) as input; however, it will be understood by a person skilled in the art that the embodiment is not limited thereto and can also be implemented in a form that generates an output using a valid document or a feature vector extracted from a valid document as input.

[0136] A processor (110) according to one embodiment can obtain an output including a label and a probability corresponding to the label by inputting at least one valid document (710) preprocessed into a plurality of different LLM-based sentiment analysis models (720). For example, the processor (110) can obtain a first result generated with relatively few resources / time by inputting at least one valid document (710) preprocessed into a first sentiment model (710) that is pre-trained to output a relatively general sentiment label for the input document. More specifically, the general sentiment label may correspond to a label that clusters precise sentiment labels corresponding to a subdivided sentiment. The processor (110) can determine a baseline sentiment for the valid document (710) based on the first result. For example, the baseline sentiment may be a sentiment corresponding to a preset number of labels among the labels included in the first result (for example, it may be selected in order of highest probability). A processor (110) can obtain a second output including a precision sentiment label corresponding to a more fine sentiment label and a corresponding probability by inputting at least one preprocessed valid document (710) and a baseline sentiment to a second sentiment model (710) that is pre-trained to output a precision sentiment label for the input document and input sentiment. For example, the precision sentiment label corresponds to an individual sentiment clustered with a general sentiment label, and may correspond to a more granular sentiment such as a precision sentiment related to corporate value, a precision sentiment related to reputation, or a precision sentiment related to risk. According to an embodiment, the processor (110) may generate a second output by inputting at least one preprocessed valid document (710) and a baseline sentiment to a second sentiment model (710) that is separately trained by category.For example, the 2-1 sentiment model (710) may be pre-trained to generate a 2-1 output related to a precision sentiment label corresponding to a precision sentiment related to corporate value, the 2-2 sentiment model (710) may be pre-trained to generate a 2-2 output related to a precision sentiment label corresponding to a precision sentiment related to reputation, and the 2-3 sentiment model (710) may be pre-trained to generate a 2-3 output related to a precision sentiment label corresponding to a precision sentiment related to risk, and the processor (110) may input at least one pre-processed valid document (710) and baseline sentiment into each of the 2-1 sentiment model to the 2-3 sentiment model to obtain a 2 output including a precision sentiment label and a corresponding probability by category. The processor (110) may generate a list of sentiment labels corresponding to the valid document (710) based on the 2 output. The list of sentiment labels may be data in which a label corresponding to the sentiment corresponding to the valid document (710) and a probability corresponding to the label are mapped.

[0137] According to an additional embodiment, a third output may be generated through a plurality of third sentiment models (710) independently trained to generate a third output including a label corresponding to at least one preprocessed valid document (710) and a probability corresponding to the label. Each of the third sentiment analysis models (710) may correspond to an LLM model trained independently of one another. The processor (110) may generate a plurality of third outputs (e.g., third-1 output, third-n output) by inputting at least one preprocessed valid document (710) into a plurality of third sentiment models (710). The processor (110) may generate a list of sentiment labels corresponding to the valid document (710) (data in which the sentiment corresponding to the valid document (710) and the probability corresponding to the sentiment are mapped) by performing a predetermined post-processing on the third output.

[0138] For example, in the post-processing process, the processor (110) may modify the labels included in the third output based on pre-set emotion set information. More specifically, the emotion set information may be information about emotions grouped into common emotions, and for example, in the emotion set information, the emotion corresponding to fear and the emotion corresponding to anxiety may be predetermined to correspond to common emotions. In this case, the processor (110) may perform post-processing in the post-processing process to re-label the label corresponding to fear and the label corresponding to anxiety in the third output as corresponding to the same label.

[0139] Additionally, in the post-processing process, the processor (110) may perform post-processing to finally determine the probability (25%) corresponding to the highest probability among the third-5 outputs as the probability corresponding to the specific label (e.g., expectation) when a specific label (e.g., a label corresponding to expectation) is included in the third-1 output (probability 10%), the third-3 output (probability 15%), and the third-5 output (25%) among the multiple third outputs.

[0140] According to one embodiment, the processor (110) can generate at least one sentiment score (740) by applying weight information included in the sentiment dictionary (730) to each label included in the sentiment label list corresponding to the valid document (710) calculated through the preceding operation.

[0141] For example, the sentiment dictionary (730) may be a database in which keywords corresponding to at least one emotion (corresponding to the individual labels described above) and weights corresponding to each keyword are mapped. The sentiment dictionary (730) generally consists of emotion keywords representing negative emotions (e.g., "shock," "pity," "disgust," etc.), emotion keywords representing neutral emotions (e.g., "annoyance," "shame," etc.), and emotion keywords representing positive emotions (e.g., "respect," "gratitude," "wonder"), and each keyword is stored with a predefined integer or real-valued weight value. An example of an implementation of the sentiment dictionary may be as illustrated in FIG. 11. For example, the keyword "doubt" may be assigned a weight of -2, "solemnity" a weight of -1, and the keyword "happiness" a weight of +2. The sentiment dictionary (730) can generally be created based on domain knowledge, results of learning from past data, and qualitative evaluations by experts, and can be used as a weight in the process of calculating a sentiment score (740) in a valid document (710) corresponding to unstructured text such as a news article. Additionally, the sentiment dictionary (730) can perform the role of assigning a large weight to a specific word when that specific word appears in the valid document (710) (e.g., rehabilitation procedure), or assigning a weight based on the frequency in past positive or negative articles.

[0142] Weights corresponding to each sentiment keyword (label) (i) included in the sentiment dictionary (730) Each label can be calculated based on mathematical formula 2, which is based on the number of times (a) derived from valid documents determined to correspond to positive events (e.g., listing, award, etc.) and the number of times (b) derived from valid documents determined to correspond to negative events (e.g., bankruptcy, sales decline, etc.) among valid documents collected at a previous point in time.

[0143]

[0144] In the example shown in Fig. 11, the weight Although it is exemplified as being normalized to the range of -2 to +2, depending on the embodiment, the weight It can be normalized into weights ranging from -1 to +1. Weights It means that the closer it is to the maximum value (e.g., +2), the more frequently it is derived from positive events, and the closer it is to the minimum value (e.g., -2), the more frequently it is derived from negative events.

[0145] A processor (110) according to one embodiment can generate at least one sentiment score (740) by reflecting a weight calculated through Equation 2 to a sentiment label list generated through the preceding operation. More specifically, the processor (110) can calculate a sentiment score (740) for each valid document (710) based on a calculation based on the probability assigned to each label included in the sentiment label list and the weight calculated through Equation 2 within the sentiment dictionary (730) for the corresponding label. For example, if a label corresponding to 'anxiety' exists on the sentiment label list calculated for the valid document (710), the processor (110) [applies] a predetermined weight to the probability determined for that label. Assign (e.g., -2), and if a label corresponding to 'joy' exists, apply a predetermined weight to the probability determined for that label. The sentiment score (740) can be calculated through an operation that assigns (e.g., +2).

[0146] For reference, the sentiment scores shown in FIG. 7 (e.g., first sentiment score to nth sentiment score) may correspond to the sentiment scores corresponding to each of the first to nth valid documents, which are individual valid documents included in at least one valid document (710).

[0147] According to an additional embodiment, the sentiment dictionary (730) may be an adaptive dictionary that is continuously updated to reflect, in addition to traditional keyword-based sentiment words, the latest social issues, industry-specific terms, and specialized terms related to finance or credit. For example, a person skilled in the art will understand that the processor may additionally include industry-specific keywords such as "restructuring," "financial difficulties," and "capital expansion" in the sentiment dictionary, and that weights may be assigned to these based on the method described above, and that these weights may be reflected in the labels obtained through the operation described above and applied to the calculation of the sentiment score. Additionally, the processor (110) may improve the precision of the sentiment score by applying a context-aware sentiment lexicon in parallel, which considers the positional relationship between keywords within a sentence or contextual nuances (e.g., double negative expressions such as "cannot help but...").

[0148] According to one embodiment, the sentiment score (740) may be provided to the user as a final labeled positive / negative / neutral based on a preset numerical standard, or may be converted into a corporate business activity score and provided by scaling it to a certain level of numerical value (e.g., 0 to 100).

[0149] The following describes a more specific method by which the processor (110) calculates the sentiment score (740).

[0150] At least one sentiment score (740) can be expressed by the following mathematical formula 3:

[0151]

[0152] Here, is among the valid documents included in at least one valid document (710). It can mean the sentiment score of the nth valid document, is defined in the sentiment dictionary (730) It may mean the weight for the th emotion keyword (i.e., the predetermined weight included in the emotion dictionary (730)), and Is In the th valid document It can refer to the probability calculated for the nth sentiment keyword (label). Probability It can be mapped together with the corresponding sentiment keyword (label) and included in the sentiment label list generated corresponding to the valid document (710) described above. The sentiment label list can be generated through the result derived from the sentiment analysis model (720) and post-processing thereof as described above.

[0153] Afterward, the processor (110) can determine a first average sentiment score (750) based on the sentiment score of each of at least one valid document.

[0154] The first average sentiment score (750) can be expressed by the following mathematical formula 4:

[0155]

[0156] Here, can mean the first average sentiment score (750), and can mean the number of valid documents, and Is It can mean the sentiment score (i.e., the normalized real value) of the i-th valid document.

[0157] The processor (110) can determine a second average sentiment score (760) based on a sentiment score determined at a preceding time prior to the time of acquiring the sentiment score of each of at least one valid document.

[0158] For example, the second average sentiment score (760) may represent the average value of previously accumulated history-based sentiment scores to be compared with the results of sentiment analysis at a specific point in time. The second average sentiment score (760) is defined as the average of sentiment scores calculated from valid documents over a past period (e.g., 180 days or 6 months) for a target company or group of companies, and may be used as a reference value to determine whether the sentiment change at the current point in time has statistical significance. That is, the second average sentiment score (760) serves as a baseline for comparison with the first average sentiment score (750) at the current point in time, and typically reflects the steady state or the sentiment flow prior to the occurrence of an event.

[0159] The second average sentiment score (760) can be expressed by the following mathematical formula 5:

[0160]

[0161] Here, can mean the second average sentiment score (760), and may refer to the number of valid documents included in a reference past period (e.g., the past 180 days), and Is It can mean the sentiment score (i.e., normalized real value) of the nth past valid document.

[0162] The processor can determine the priority of at least one valid document based on a first average sentiment score (750), a second average sentiment score (760), and a standard deviation determined based on the sentiment score determined at a preceding time.

[0163] For example, the priority in this specification may refer to the priority of at least one valid document for generating summary information provided to a user, based on the premise that at least one valid document (i.e., a plurality of valid documents) is grouped according to a specific date. Specifically, the priority of at least one valid document at a first time point and the priority of at least one valid document at a second time point may be different from or the same. For example, if the priority of at least one valid document at a first time point is higher than the priority of at least one valid document at a second time point, the processor may provide summary information based on at least one valid document at a first time point to the user preferentially over summary information based on at least one valid document at a second time point.

[0164] A detailed description of the operation for determining the priority of at least one valid document determined at a specific point in time (e.g., a first point in time or a second point in time) is provided below.

[0165] Specifically, the processor can determine the Z-score (770) of at least one valid document based on the first average sentiment score (750), the second average sentiment score (760), and the standard deviation.

[0166] For example, the Z-score (770) may represent an indicator that quantitatively indicates the degree of anomaly or deviation from a standard criterion of the sentiment score calculated at the current point in time. The Z-score (770) can be calculated by normalizing the difference between the first average sentiment score (750) and the past second average sentiment score (760) by the standard deviation of the past sentiment scores. That is, the Z-score (770) indicates how drastically the current sentiment level has changed compared to the past average, and can serve as a measure to detect the management status of the target company or the possibility of external risks occurring at an early stage. For example, since the greater the absolute value of the Z-score (770), the greater the deviation from the past pattern, the processor may classify valid documents with a large absolute value of the Z-score (770) as the first priority for review.

[0167] The Z-score (770) can be expressed by the following mathematical formula 6:

[0168]

[0169] Here, can mean Z-score (770), and can mean the first average sentiment score (750), and can mean the second average sentiment score (760), and can mean the standard deviation corresponding to the second average sentiment score (760).

[0170] The processor (110) can determine the priority of at least one valid document as the first priority for priority review if the Z-score (770) is not included in a predetermined score range, and determine the priority of at least one valid document as the second priority for general review if the Z-score (770) is included in a predetermined score range.

[0171] A processor (110) according to one embodiment can provide a means of responding to an abnormal situation by calculating a Z-score (770) that reflects a time-series change situation by repeatedly calculating a first average sentiment score (750) and a second average sentiment score (760) based on the current time point.

[0172] The priority of at least one valid document can be expressed by the following mathematical formula 7:

[0173]

[0174] Here, can mean the priority of at least one valid document, and can mean Z-score (770), and is the threshold score of a predetermined score range, It is the first priority, can mean second priority.

[0175] For example, if at a first time point the priority of at least one valid document is a first priority regarding priority review and at a second time point the priority of at least one valid document is a second priority regarding general review, the processor may provide summary information based on at least one valid document at the first time point to the user preferentially over summary information based on at least one valid document at the second time point.

[0176] In contrast, the processor (110) may provide summary information based on at least one valid document at the second time point to the user in priority over summary information based on at least one valid document at the first time point, when the priority of at least one valid document at the first time point is a second priority regarding general review and the priority of at least one valid document at the second time point is a first priority regarding priority review.

[0177] Additionally, the processor (110) can provide summary information based on at least one valid document at the time that precedes the first time or the second time, if the priority of at least one valid document at the first time is a second priority regarding general review and the priority of at least one valid document at the second time is a second priority regarding general review.

[0178] For example, the first priority regarding priority review may refer to a priority grade that designates a valid document as a target for rapid and preemptive analysis when the sentiment score of a valid document related to a target company deviates significantly from statistical criteria. Specifically, if the Z-score (770) of a valid document falls outside a predefined score range (e.g., greater than or less than ±1.5), the valid document is considered to have undergone an abnormal sentiment change and is judged to be a document with a high probability of risk or major event occurring for the target company. The processor (110) may classify the valid document as first priority and process it so that it is immediately provided to the user through a warning or notification system, or is passed on to a subsequent analysis stage (e.g., LLM summary, risk analysis, etc.).

[0179] For example, the second priority regarding general review may refer to a normal priority grade in which valid documents are processed according to standard analysis procedures when the sentiment score of a valid document does not deviate statistically significantly from the historical average. The second priority refers to the case where the Z-score (770) falls within a preset threshold score range and corresponds to a state where it is determined that there are no major abnormal changes in corporate-related news. In this case, the processor classifies the valid document as a low-priority document and can include it in general workflows such as generating periodic reports or processing lower-priority summaries.

[0180] For example, the score range is a range that defines the criteria for determining whether a document is considered an abnormal sign when the sentiment analysis-based Z-score (770) is above a certain level, and includes a preset threshold. Here, the score range may be dynamically set or adjusted by the user, taking into account the size of the company, industry characteristics, or past sentiment variability. For example, the processor (110) may provide a function to provide notifications to the company based on positive events when the Z-score (770) exceeds a positive threshold, and conversely, may provide a function to provide notifications to the company based on negative events when the Z-score (770) is below a negative threshold.

[0181] The processor (110) can provide a means for the company corresponding to the user to recognize an abnormal situation by providing an alarm to the user for a point in time (date) corresponding to at least one valid document determined as the first priority, thereby providing a notification that an abnormal situation has occurred on a specific date. In addition, if the criterion for determining the first priority is below a negative threshold, the processor (110) can provide a preset number of valid documents in order of lowest sentiment score among the valid documents corresponding to that date, or generate and provide a summary of the documents through LLM. Conversely, if the criterion for determining the first priority exceeds a positive threshold, the processor (110) can provide a preset number of valid documents in order of highest sentiment score among the valid documents corresponding to that date, or generate and provide a summary of the documents through LLM.

[0182] FIG. 8 shows an example of a method for expressing an emotional score in an information providing device according to one embodiment.

[0183] FIG. 8 is a diagram illustrating an example in which a processor (e.g., the processor (110) of FIG. 1) according to one embodiment of the present invention calculates a sentiment score based on valid documents related to a company and organizes it into a table form to provide it visually.

[0184] Specifically, FIG. 8 illustrates documents determined to contain negative content through sentiment analysis among a number of valid documents collected for a target company, and the sentiment scores corresponding to each valid document. In the example illustrated in FIG. 8, a total of three valid documents were extracted for one company indicated as "Target Company," and a summary phrase for each document and the sentiment score calculated from that document are shown in a paired form.

[0185] The sentiment score is a value calculated by a sentiment analysis model based on labels derived from the sentiment analysis of each valid document and the probability values ​​corresponding to those labels, and is typically expressed in the form of a real number. For example, in Figure 8, sentiment scores are shown as -1.625, -0.751, -1.175, etc., and the sentiment score quantitatively indicates how much negative content each valid document contains. A lower sentiment score (a larger negative score) indicates that the document is more negative.

[0186] The content illustrated in FIG. 8 can contribute to identifying management risks of a target company in advance and intuitively grasping signs of problems by providing quantitative sentiment analysis results for collected valid documents (e.g., news documents). In particular, the processor can provide information that enables rapid detection of corporate situation volatility and crisis signals through the comparison and analysis of sentiment scores for multiple documents.

[0188] FIG. 9 shows an example of visualizing a change in sentiment score based on the time of a target company's application for rehabilitation in an information providing device according to one embodiment.

[0189] FIG. 9 is a diagram showing an example of visualizing changes in sentiment-based analysis results in a time series form based on the time of the target company's application for rehabilitation in a processor according to one embodiment of the present invention (e.g., processor (110) of FIG. 1).

[0190] As illustrated in Fig. 9, publication dates (e.g., news article publication dates) are arranged in chronological order in the vertical direction, and for each date, quantitative indicators such as the number of articles, the average sentiment score (or sentiment score) of the articles collected on that date, the sentiment Z-score, and the article Z-score are presented in parallel. For example, an event called a "reorganization application" is indicated as having occurred as of March 4, 2025, and rapid changes in the sentiment score and related statistics for the target company in the period before and after this point are illustrated.

[0191] Specifically, immediately before the event occurs, the number of articles and the sentiment score show a relatively stable trend, but immediately after the event occurs, the number of articles explodes and the sentiment score drops sharply, while the sentiment Z-score and article Z-score show extreme values ​​and exhibit a rapidly changing pattern. Figure 9 illustrates an example in which a processor can perform the technical function of an early warning system for a target company by visually demonstrating that quantitative indicators based on sentiment analysis can show significant abnormal signs from a point in time prior to the occurrence of a corporate risk event.

[0193] FIG. 10 shows an example of a user interface that outputs summary information provided based on LLM output and priority in an information providing device according to one embodiment.

[0194] FIG. 10 illustrates an example of summary information provided based on the output of a large language model and the priority of valid documents in a processor according to one embodiment of the present invention (e.g., the processor (100) of FIG. 1).

[0195] For example, the summary information illustrated in Fig. 10 can provide information that allows users to quickly and accurately understand the risk situation of the target company by integrating article summary results generated through a large language model, quantified sentiment scores, and meta-information that can be used to determine whether to prioritize review.

[0196] Specifically, the processor determines the priority of at least one valid document, applies at least one valid document to a large language model to generate at least one LLM output and priority, displays it in a binary classification form (1: negative or 0: not applicable), and provides relative evaluation results, such as the bottom 20%, based on the relative distribution of sentiment scores through weighted operations based on a sentiment dictionary, along with summary information of the target company.

[0197] In addition, the processor utilizes LLM output to provide summaries extracted from news headlines and body content (i.e., content contained in at least one valid document), and can provide key sentences regarding circumstances of a company's deteriorating business conditions (e.g., insolvency, failure to settle, distribution crisis, etc.) identified in actual news articles. Furthermore, the processor provides various meta-information, such as the publication date of the news (i.e., the publication date of the document), media outlet information, the number of similar articles published, and the scope of interest coverage, thereby providing credibility and attention to the information.

[0198] According to the present embodiment, the processor can calculate a log-weighted moving average (LWMA) and a variance-based outlier indicator together to precisely evaluate the time-series trend of the sentiment score.

[0199] For example, the processor can calculate the log-weighted moving average for a set of sentiment scores {s1, s2, ... , sn} included in period T as shown in the following mathematical formula 8.

[0200]

[0201] Here, represents the exponential weight for the most recent point in time.

[0202] In addition, the processor [represents] the log variance value for the above set It can be calculated as shown in the following mathematical formula 9.

[0203]

[0204] The processor is the current sentiment score For this, the following anomaly detection indicator A can be calculated through mathematical formula 10.

[0205]

[0206] The processor has a threshold value for A. If it exceeds [value], the document may be determined to be a sharp sentiment outlier and classified as a priority review subject. This provides a means for risk detection that reflects more complex volatility compared to a simple Z-score.

[0207] According to another embodiment, the processor can produce an index including a polynomial transformation and a normalization term to evaluate the complex correlation between the sentiment vector and the topic vector.

[0208] More specifically, for a specific document dj, the processor can calculate the joint indicator Cj as in Equation 10, where the sentiment vector is vs=(s1,s2, ... ,sm) and the topic vector is vt=(t1,t2, ... ,tm).

[0209]

[0210] Here, is a very small positive number for the stability of the denominator, and α is the weight of the two terms (0 ≤ α ≤ 1).

[0211] In the above equation, the first term emphasizes correlation by squaring the normalized inner product of the sentiment vector and the topic vector, and the second term non-linearly reflects the weighted sum between the sentiment score and the topic item by log-transforming it.

[0212] If Cj is above a predefined threshold γ, the processor can determine that the document is a key issue closely related to corporate risk and immediately trigger an alert state.

[0213] The device described above may be implemented as a hardware component, a software component, and / or a combination of a hardware component and a software component. For example, the device and components described in the embodiments may be implemented using one or more general-purpose or special-purpose computers, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing unit may execute an operating system (OS) and one or more software applications executed on said operating system. Additionally, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing unit may be described as being used as a single unit, but those skilled in the art will understand that the processing unit may include a plurality of processing elements and / or a plurality of types of processing elements. For example, the processing unit may include multiple processors or one processor and one controller. Additionally, other processing configurations, such as parallel processors, are also possible.

[0214] Software may include computer programs, code, instructions, or a combination of one or more of these, and may configure a processing unit to operate as desired or command the processing unit independently or collectively. Software and / or data may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave so as to be interpreted by the processing unit or to provide instructions or data to the processing unit. Software may be distributed over networked computer systems and may be stored or executed in a distributed manner. Software and data may be stored on one or more computer-readable recording media.

[0215] The method according to the embodiment may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program instructions, data files, data structures, etc., either alone or in combination. The program instructions recorded on the medium may be those specifically designed and configured for the embodiment, or they may be those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc. The hardware devices described above may be configured to operate as one or more software modules to perform the operation of the embodiment, and vice versa.

[0216] Although the embodiments have been described above with reference to limited examples and drawings, those skilled in the art can make various modifications and variations from the description above. For example, suitable results can be achieved even if the described techniques are performed in a different order than described, and / or the components of the described system, structure, device, circuit, etc. are combined or assembled in a form different from described, or replaced or substituted by other components or equivalents.

[0217] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims set forth below.

Claims

Claim 1 In an information providing device through sentiment analysis based on a large-scale language model, at least one processor; and at least one memory operatively connected to the at least one processor; wherein, when executed, the at least one memory stores instructions that cause the at least one processor to obtain a corpus regarding the target enterprise through identification information of the target enterprise included in the list of managed operators based on receiving a list of managed operators, and to extract at least one valid document containing a target keyword regarding the business operation of the target enterprise from the at least one document based on performing preprocessing on at least one document included in the corpus, and to apply the at least one valid document to a sentiment analysis model to obtain a sentiment score for each of the at least one valid document, determine the priority of the at least one valid document based on the sentiment score of each of the at least one valid document, and to provide summary information of the target enterprise based on at least one LLM output generated by applying the at least one valid document to a large language model (LLM) and the priority, and wherein, when executed, the at least one memory stores instructions that cause the at least one processor to... An information providing device that stores instructions causing to determine a priority of at least one valid document based on the first average sentiment score determined based on each of the above sentiment scores, a second average sentiment score determined based on the sentiment score determined at a preceding time prior to the time of acquiring the sentiment score of each of the at least one valid document, and a standard deviation determined based on the first average sentiment score, the second average sentiment score, and the sentiment score determined at the preceding time prior. Claim 2 In claim 1, the at least one memory is an information providing device that, at execution, stores instructions that cause the at least one processor to determine the identification information including at least one of the target company’s Korean name, English name, English abbreviation, stock code, standard code, alias, representative name, or any combination thereof, generate a search keyword based on at least one of the Korean name, English name, English abbreviation, stock code, standard code, abbreviation, representative name, or any combination thereof, and apply the search keyword to an information providing platform regarding a news database, a digital content repository, or an online posting system to acquire the corpus. Claim 3 In claim 1, the at least one memory stores instructions that, at execution, cause the at least one processor to extract a first comparison document and a second comparison document from the at least one document included in the corpus, apply the first comparison document and the second comparison document to a sentence transformer (SBERT) model learned to calculate semantic similarity between sentences, obtain a document score regarding the content, structure, and form similarity of the first comparison document and the second comparison document, and remove one of the first comparison document or the second comparison document from the corpus based on the document score exceeding a predetermined threshold score, thereby causing preprocessing to be performed on the at least one document included in the corpus. Claim 4 In claim 1, the above-mentioned at least one memory stores instructions that, at execution, cause the at least one processor to determine a first keyword including at least one of the Korean name of the target company, the English name of the target company, or any combination thereof, determine a second keyword for extracting a document created from a predetermined information providing platform from the at least one document, apply a third keyword predetermined as unrelated to the operation of the target company to the at least one document to remove a document containing the third keyword from the at least one document, and apply the first keyword and the second keyword to the at least one document from which the document containing the third keyword has been removed to extract the at least one valid document. Claim 5 In claim 1, the information providing device comprises, wherein, at least one memory, when executed, at least one processor obtains a sentiment score for each of the at least one valid document by applying the at least one valid document to the sentiment analysis model, and the sentiment score is normalized to a value included in a predetermined real number range and includes a criterion for classifying each of the at least one valid document as a negative document or a positive document. Claim 6 In claim 5, the above-mentioned at least one memory stores instructions that, at execution, cause the above-mentioned at least one processor to preprocess each of the above-mentioned at least one valid document to extract feature vectors, apply the feature vectors to the above-mentioned sentiment analysis model which is composed of a structure in which a plurality of layers are stacked and includes a Multi-head Attention and a Position-wise Feed Forward Network in the plurality of layers, obtain the sentiment score based on the attention operation, and classify the sentiment score to estimate at least one of the positive document or negative document of each of the above-mentioned at least one valid document. Claim 7 delete Claim 8 In claim 1, the at least one memory stores instructions that, at execution, cause the at least one processor to determine the Z-score of the at least one valid document based on the first average sentiment score, the second average sentiment score, and the standard deviation, and if the Z-score is not included in a predetermined score range, determine the priority of the at least one valid document as a first priority for priority review, and if the Z-score is included in a predetermined score range, determine the priority of the at least one valid document as a second priority for general review. Claim 9 In claim 1, the information providing device stores commands that cause the LLM output to include at least one of the summary keywords and sentiment scores included in the at least one valid document according to a predetermined prompt. Claim 10 In claim 1, the at least one memory stores instructions that cause the at least one processor to identify the past identification information, based on a change history between the past identification information corresponding to the target enterprise and the past identifier of the target enterprise and the current identification information corresponding to the current identifier of the target enterprise, so as to include the at least one document corresponding to the past identification information in the corpus. Claim 11 A method for providing information through large-scale language model-based sentiment analysis performed by an information providing device, comprising: an operation of obtaining a corpus regarding a target company through identification information of the target company included in the list of management business operators based on receiving a list of management business operators; an operation of extracting at least one valid document containing a target keyword regarding the business operation of the target company from the at least one document based on performing preprocessing on at least one document included in the corpus; an operation of applying the at least one valid document to a sentiment analysis model to obtain a sentiment score for each of the at least one valid document; and an operation of determining the priority of the at least one valid document based on the sentiment score of each of the at least one valid document. A method for providing information, comprising: an operation of providing summary information of a target company based on at least one LLM output generated by applying at least one valid document to a large language model (LLM) and the priority; an operation of determining a first average sentiment score based on the sentiment score of each of the at least one valid document; an operation of determining a second average sentiment score based on the sentiment score determined at a preceding time point prior to the time of acquiring the sentiment score of each of the at least one valid document; and an operation of determining the priority of the at least one valid document based on the first average sentiment score, the second average sentiment score, and the standard deviation determined based on the sentiment score determined at the preceding time point.