Method and system utilizing large language model for topic allocation based on quotations

The quotation-based topic assignment method using a large language model addresses the limitations of traditional methods by extracting and assigning topics at the quotation level, enhancing accuracy and reliability in topic analysis.

WO2026111366A1PCT designated stage Publication Date: 2026-05-28LG MANAGEMENT DEV INST CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/019062
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-11-18
Filing Date
2025-11-18
Publication Date
2026-05-28

AI Technical Summary

Technical Problem

Existing topic assignment methods struggle to provide precise evidence at the span level in long documents or paragraphs with complex topics, leading to ambiguous semantic boundaries and low interpretability, especially in multi-topic corpora, where traditional methods fail to clearly separate document clusters by topic and identify fine-grained topic structures.

Method used

A quotation-based topic assignment method using a large language model that extracts and assigns topics at the quotation level within content, utilizing a large language model to specify topics and extract semantically consecutive word sets as quotations, providing clear topic analysis results.

Benefits of technology

Enhances the accuracy and reliability of topic analysis by explicitly deriving sentences or sentence segments as topic bases, resolving topic contamination and improving semantic consistency and document comprehension.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025019062_28052026_PF_FP_ABST
    Figure KR2025019062_28052026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a method and a system utilizing a large language model for topic allocation based on quotations. According to the present invention, a computer-implemented method utilizing a large language model for topic allocation based on quotations may comprise the steps of: specifying content to be analyzed, which includes at least one sentence; specifying at least one topic related to the at least one sentence included in the content to be analyzed; extracting, by using a large language model, at least one quotation corresponding to each of the at least one topic from the content to be analyzed; and providing, by using the topic and the quotation, a topic analysis result for the content to be analyzed.
Need to check novelty before this filing date? Find Prior Art

Description

Quote-based topic assignment method and system using a large language model

[0001] The present invention relates to a quotation-based topic assignment method and system using a large language model, which assigns topics to quotations constituting at least one sentence using a large language model.

[0002] The dictionary definition of artificial intelligence is a technology that realizes human learning, reasoning, perception, and natural language understanding abilities through computer programs. This artificial intelligence has achieved rapid development through deep learning.

[0003] In particular, driven by the advancement of artificial intelligence, various language models have been developed. These models have reached a level where they not only recognize text and understand its meaning but also extract and classify information from vast amounts of text-based data, such as documents, and even generate text directly.

[0004] These language models are actively utilized in various fields and exist in diverse areas where text-based tasks can be performed, such as search engines, document creation (e.g., resume writing, report writing, posting, etc.), free conversation on various topics, data parsing from given text (e.g., data summarization, classification, etc.), provision of expertise, programming, and converting given sentences into sentences of an appropriate style.

[0005] In this regard, Large Language Models (LLMs) have recently emerged, which understand and generate human language through prior training on vast amounts of text data. Unlike traditional chatbots, which are manually built and provide only limited answers, LLMs are demonstrating innovation in the artificial intelligence market by showcasing technological capabilities that allow them to communicate naturally, almost like humans, and provide fast and accurate information.

[0006] Recently, research is actively underway to utilize large language models in topic assignment technology, which structures 'what constitutes what topic' within large-scale corpora.

[0007] Specifically, traditional topic models used in the field of existing topic assignment techniques (e.g., LDA-based document-word distribution modeling, embedding-clustering-based BERTopic, etc.) tend to estimate the topic of a document using word-level probability distributions or distances between document embeddings. These existing topic assignment methods have the problem that it is difficult to provide precise evidence at the span level in long documents or paragraphs with complex topics, and the interpretability of the results is limited.

[0008] Furthermore, traditional topic models had limitations in terms of topic purity, as contamination problems arose where different semantic domains were mixed within a single topic, leading to low consistency among documents within the same topic.

[0009] Furthermore, existing Document-Based Topic Assignment (DBTA) methods are limited to assigning a single representative topic to each document or distributing topics across the entire document; consequently, there was a limitation in that it was difficult to specifically identify which part corresponds to which topic at the sentence or segment level, even when multiple topics coexisted within a document.

[0010] In this structure, since topics are expressed only at the “document level” or “word distribution,” users must go through the process of inferring the subject indirectly by examining top keyword lists or document sets for each topic; furthermore, because it is difficult to verify direct evidence sentences for a specific topic, there are limitations in ensuring the interpretability and reliability of the results.

[0011] In particular, in a multi-topic corpus environment where a single document contains both a main topic and multiple subtopics, this existing document-based topic assignment method represents the entire document as a single topic or a mixture of a few topics, which leads to ambiguous semantic boundaries between topics and frequent instances of topic clusters penetrating each other.

[0012] As a result, existing document-based topic assignment methods fail to clearly separate document clusters by topic, and there is a problem in that it is difficult to identify a fine-grained topic structure using only the sets of keywords and documents extracted by topic.

[0013] To address these issues, there is a need for topic assignment technology that utilizes large language models to assign topics at the quotation level within the content under analysis.

[0014] The present invention aims to provide a quotation-based topic assignment method and system using a large language model, which can assign topics in units of quotations in content to be analyzed using a large language model.

[0015] Specifically, the present invention is to provide a quotation-based topic assignment method and system using a large language model, which can extract at least one quotation corresponding to each of at least one topic from content to be analyzed using a large language model.

[0016] Furthermore, the present invention aims to provide a quotation-based topic assignment method and system using a large language model capable of providing topic analysis results for content to be analyzed by utilizing topics and quotations.

[0017] To solve the problem described above, a quotation-based topic assignment method using a large language model according to the present invention, performed by a computer, may include the steps of: specifying content to be analyzed consisting of at least one sentence; specifying at least one topic related to the at least one sentence constituting the content to be analyzed; using a large language model to extract at least one quotation corresponding to each of the at least one topic from the content to be analyzed; and using the topic and the quotation to provide a topic analysis result for the content to be analyzed.

[0018] In one embodiment, the step of specifying the at least one topic may include: generating a first prompt requesting the generation of at least one topic candidate composed of at least one of a domain and a category based on the content to be analyzed; processing the first prompt as input to the large language model to obtain the at least one topic candidate from the large language model; and specifying the at least one topic associated with the at least one sentence based on the at least one topic candidate.

[0019] In one embodiment, the step of extracting at least one quotation may include generating a second prompt requesting the extraction of at least one quotation based on the content to be analyzed and the at least one topic candidate, and processing the second prompt as input to the large language model to extract at least one quotation from the content to be analyzed from the large language model.

[0020] In one embodiment, the second prompt may be configured so that the large language model extracts at least one word set consisting of at least one word that is semantically consecutive in the at least one sentence as the at least one quotation.

[0021] In one embodiment, the method may further include the step of assigning a specific topic among the at least one topic to each of the at least one quotation using the large language model.

[0022] In one embodiment, the second prompt may be configured to map the specific topic among the at least one topic to at least one specific quotation corresponding to the specific topic.

[0023] In one embodiment, the step of assigning the specific topic may include the step of processing the second prompt as input to the large language model and the step of the large language model outputting at least one topic assignment information in which the specific topic and the at least one specific quotation are mapped.

[0024] In one embodiment, the topic analysis result may be configured such that, based on the at least one topic assignment information, the at least one quotation included in the content to be analyzed is classified and displayed according to the at least one topic.

[0025] In one embodiment, the topic analysis result may include quotation reference information related to the location of the at least one quotation included in the at least one topic assignment information in the content to be analyzed.

[0026] In one embodiment, the method may further include the step of generating topic analysis content from the content to be analyzed according to a user query received from a user terminal, and the step of providing the generated topic analysis content to the user terminal as a response to the user query.

[0027] In one embodiment, the topic analysis content may be configured to include at least one of the topic analysis result corresponding to the user query and the response information to the user query.

[0028] In one embodiment, the step of generating the topic analysis content may include the step of generating a prompt requesting the generation of the topic analysis content based on the user query and the content to be analyzed, and the step of inputting the prompt into the large language model to obtain the topic analysis content related to the user query from the large language model.

[0029] In one embodiment, the step of generating the prompt may involve collecting user information corresponding to a user account logged into the user terminal, and generating the prompt that requests the large language model to generate the topic analysis content based on the user information and the user query.

[0030] In one embodiment, the response information may be configured to include summary information related to the user query extracted from the topic analysis result.

[0031] In one embodiment, the method may further include the step of determining a main topic and a sub-topic among the at least one topic related to the content to be analyzed.

[0032] In one embodiment, in the step of determining the main topic and subtopic, the main topic among the at least one topic can be determined based on the number of at least one citation mapped to each of the at least one topic.

[0033] In one embodiment, the topic analysis result may be configured such that main topic assignment information corresponding to the main topic and sub-topic assignment information corresponding to the sub-topic are distinguished.

[0034] In one embodiment, the quotation may be a set of at least one word consisting of at least one word that is semantically consecutive in the at least one sentence.

[0035] Meanwhile, the system includes a memory for storing instructions and at least one processor electrically connected to the memory, and when the instructions are executed by the at least one processor, the at least one processor identifies an analysis target content consisting of at least one sentence, identifies at least one topic related to the at least one sentence constituting the analysis target content, extracts at least one quotation corresponding to each of the at least one topic from the analysis target content using a large language model, and provides a topic analysis result for the analysis target content using the topic and the quotation.

[0036] A program according to the present invention is a program that is executed by one or more processes in an electronic device and is stored in a computer-readable recording medium, and may include instructions for performing the steps of: specifying content to be analyzed composed of at least one sentence; specifying at least one topic related to the at least one sentence constituting the content to be analyzed; extracting at least one quotation corresponding to each of the at least one topic from the content to be analyzed using a large language model; and providing a topic analysis result for the content to be analyzed using the topic and the quotation.

[0037] As described above, the quotation-based topic assignment method and system using a large language model according to the present invention can accurately assign topics to multiple sentences constituting the content to be analyzed based on quotations by utilizing a large language model. Through this, the accuracy and reliability of the topic analysis results can be improved by explicitly deriving the sentences or sentence segments that serve as the basis for topics in various forms of content (e.g., documents, articles, reviews, etc.).

[0038] Furthermore, the quotation-based topic assignment method and system using a large language model according to the present invention utilizes a large language model to extract quotations corresponding to at least one topic from the content to be analyzed and to assign topics on a quotation basis. Through this, the problem of topic contamination can be resolved when analyzing topics of the content to be analyzed, and semantic consistency within a single topic can be enhanced.

[0039] Furthermore, the quotation-based topic assignment method and system using a large language model according to the present invention can provide a topic analysis result including topic assignment information for content to be analyzed by combining topics and quotations. Through this, users can grasp the core topic of the content to be analyzed, the relationships between topics, and key quotations corresponding to each topic at a glance, and consequently, improve document comprehension and decision-making efficiency.

[0040] FIG. 1 illustrates an example of a block diagram of a computing system in which the present invention can be implemented.

[0041] FIG. 2 illustrates an example of a block diagram of a computing device that may be included in a user computing device, a server computing system, and a training computing system, as an embodiment of a computing system in which the present invention can be implemented.

[0042] Figure 3 illustrates an example of a block diagram from another perspective of a computing device, which is one of the components of a computing system.

[0043] FIG. 4 is a conceptual diagram illustrating a quotation-based topic assignment system using a large language model according to the present invention.

[0044] FIG. 5 is a conceptual diagram for explaining the concept related to quotation-based topic assignment according to the present invention.

[0045] FIG. 6 is a flowchart illustrating a quotation-based topic assignment method using a large language model according to the present invention.

[0046] FIGS. 7, FIGS. 8a, FIGS. 8b, FIGS. 9, and FIGS. 10 are conceptual diagrams for specifically explaining a quotation-based topic assignment method using a large language model according to the present invention.

[0047] FIGS. 11a to 11d are conceptual diagrams illustrating various embodiments using a quotation-based topic assignment method using a large language model according to the present invention.

[0048] Hereinafter, embodiments disclosed in this specification will be described in detail with reference to the attached drawings. Identical or similar components are assigned the same reference number regardless of the drawing symbols, and redundant descriptions thereof will be omitted. The suffixes "module" and "part" used for components in the following description are assigned or used interchangeably solely for the ease of drafting the specification and do not have distinct meanings or roles in themselves. Furthermore, in describing the embodiments disclosed in this specification, if it is determined that a detailed description of related prior art could obscure the essence of the embodiments disclosed in this specification, such detailed description will be omitted. Additionally, the attached drawings are intended only to facilitate understanding of the embodiments disclosed in this specification; the technical concept disclosed in this specification is not limited by the attached drawings, and it should be understood that they include all modifications, equivalents, and substitutions that fall within the spirit and technical scope of the present invention.

[0049] Terms including ordinal numbers, such as first, second, etc., may be used to describe various components, but said components are not limited by said terms. These terms are used solely for the purpose of distinguishing one component from another.

[0050] When it is stated that one component is "connected" or "connected" to another component, it should be understood that while it may be directly connected or connected to that other component, there may also be other components in between. On the other hand, when it is stated that one component is "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between.

[0051] A singular expression includes a plural expression unless the context clearly indicates otherwise.

[0052] In this application, terms such as “comprising” or “having” are intended to specify the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0053] Hereinafter, the present invention will be examined in more detail with reference to the attached drawings. FIG. 1 illustrates an example of a block diagram of a computing system in which the present invention can be implemented. FIG. 2 illustrates an example of a block diagram of a computing device that may be included in a user computing device, a server computing system, and a training computing system as an embodiment of a computing system in which the present invention can be implemented. FIG. 3 illustrates an example of a block diagram of a computing device in another aspect that is one of the components of a computing system.

[0054] Meanwhile, FIG. 1 illustrates an example of a block diagram of a computing system in which the present invention can be implemented. In this regard, a quotation-based topic assignment system using a large language model according to the present invention can be implemented through a computing device described below and can perform data processing related to the quotation-based topic assignment method using the large language model described above.

[0055] Referring to FIG. 1, a computing system (10000) that performs a quotation-based topic assignment method using a large language model of the present invention may include at least one computing device. At this time, the at least one computing device may be a single processor or a multiprocessor computing device.

[0056] The components of at least one computing device of the present invention may include various hardware components such as one or more processors, memory, other hardware, and a system bus (not shown) that connects various system components so that they can transmit and receive data to and from each other (e.g., telecommutatively connected, physically connected, electrically connected), and the components of at least one computing device are not limited thereto and may be very diverse.

[0057] Meanwhile, at least one computing device included in a computing system (10000) that performs a quotation-based topic assignment method using a large language model may be connected to communicate via a network (1070). For example, at least one computing device included in the computing system (10000) may be clustered or part of a local area network (LAN). Additionally, at least one computing device may be part of a wide area network (WAN) or connected to at least one of a client-server network and a peer-to-peer network within the cloud.

[0058] Meanwhile, when at least one computing device is used in at least one of a network environment and a cloud computing environment, the at least one computing device may be connected to at least one of a public and private network through a network interface or adapter. In one embodiment, other communication connection devices, such as a modem, may be used to establish communication through the network. The modem may be at least one of an internal modem and an external modem, and may be connected to a system bus through a network interface or a specific mechanism, etc. A wireless network component consisting of an interface and an antenna may be coupled to the network through a device such as an access point, a peer computer, etc. In the present invention, the method of connecting at least one computing device to communicate through the network (1070) is not limited, and it may be connected to communicate in a manner different from the described example.

[0059] Furthermore, other computer-type devices and / or systems not shown in FIG. 1 may also interact technically with at least one computing device or other system through one or more connections to the network (1070) via a network interface. Here, the network interface may include network interface equipment such as a physical network interface controller (NIC) or a virtual network interface (VIF).

[0060] The network (1070) of the present invention may include various forms such as the Internet, WLAN (Wireless LAN), Wi-Fi (Wireless-Fidelity), Wi-Fi (Wireless Fidelity) Direct, DLNA (Digital Living Network Alliance), WiBro (Wireless Broadband), WiMAX (World Interoperability for Microwave Access), HSDPA (High Speed ​​Downlink Packet Access), HSUPA (High Speed ​​Uplink Packet Access), LTE (Long Term Evolution), LTE-A (Long Term Evolution-Advanced), 5G (5th Generation Mobile Telecommunication), Bluetooth (Bluetooth™ Frequency Identification), Infrared Data Association (IrDA), UWB (Ultra-Wideband), ZigBee, NFC (Near Field Communication), Wireless USB (Wireless Universal Serial Bus), etc., and in the present invention, data transmission may be performed based on standard communication protocols such as TCP / IP, HTTP, SSL, etc.

[0061] A computing system (10000) that performs a quotation-based topic assignment method using a large language model according to the present invention may include at least one of a user computing device (1010), a training computing system (1050), and a server computing system (1030).

[0062] A user computing device (1010) according to the present invention may be understood as a computing device comprising at least one processor (1011) and memory (1012) that perform a quotation-based topic assignment method using a large language model. For example, the user computing device (1010) may include at least one computing device among a smartphone, a smart TV, a laptop computer, a desktop computer, a digital broadcasting terminal, a PDA (personal digital assistants), a PMP (portable multimedia player), a navigation device, a slate PC, a tablet PC, an ultrabook, a wearable device (e.g., a smartwatch, a smart glass, and a head-mounted display).

[0063] At least one processor (1011) constituting the user computing device (1010) may include one or more general-purpose processors and / or one or more special-purpose processors. For example, at least one processor (1011) constituting the user computing device (1010) may be composed of at least one of a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), a neural network processing unit (NPU), an arithmetic logic unit (ALU), a floating-point arithmetic unit (FPU), an application integrated circuit, an application semiconductor (ASIC), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, and / or electrical units for performing other functions, or a plurality of electrically connected processors.

[0064] Furthermore, at least one processor (1011) may be configured to execute computer-readable instructions contained in memory (1012) and / or other instructions described herein.

[0065] The memory (1012) constituting the user computing device (1010) according to the present invention may include volatile memory, non-volatile memory, fixed media, removable media, magnetic media, optical media, semiconductor media and / or other types of physically durable storage media.

[0066] For example, memory (1012) may include one or more non-transient / transient computer-readable storage media such as RAM, ROM, HDD (Hard Disk Drive), SSD (Solid State Disk), SSD (Silicon Disk Drive), EEPROM, EPROM, flash memory device, magnetic disk, etc., and may include web storage of a server that performs memory storage functions on the internet. Such memory (1012) may store data and commands necessary for the operation of an application to generate topic analysis results for the content to be analyzed by the at least one processor (1011) using a large language model to extract quotes from the content to be analyzed and assign topics to the quotes.

[0067] A user computing device (1010) may include one or more user input components (1021) that detect user input. For example, the user input component (1021) may also be referred to as a user interface module. The user input component (1021) may include a touch screen, a computer mouse, a keyboard, a keypad, a touchpad, a trackball, a joystick, a voice recognition module, or other similar devices. However, the present invention does not limit the type of user input component (1021). In this case, the user input component (1021) in the present invention does not necessarily mean a hardware means, but can be understood as a channel for receiving input from a user. Meanwhile, the user of the present invention may refer to an automated agent, script, playback software, etc., that operates on behalf of one or more people.

[0068] A user can interact with a computing system (10000) including at least one computing device through input text, touch, voice, movement, computer vision, gestures and / or other forms of input / output using a user input component (1021). For example, the user input component (1021) may include one or more of a command line interface (CLI), a graphical user interface (GUI), a natural user interface (NUI), a voice command interface and / or other user interface (UI) representations.

[0069] Between the user input component (1021) and the user computing device (1010), one or more application programming interface (API) calls may be made based on user input received from the user interface and / or network.

[0070] Here, the expression "based on" may be interpreted to include cases where it is based on the use of a specific configuration, modified from, derived from, influenced by, dependent on, or otherwise derived from a specific configuration. In some embodiments, an API call may be configured for a specific API, which may be interpreted or converted into an API call configured for another API. Here, an API may refer to a defined interface or connection between computers or between computer programs.

[0071] In one embodiment, the user computing device (1010) may store at least one machine learning model (1020). For example, the user computing device (1010) may be various machine learning models, such as a plurality of neural networks (e.g., deep neural networks) that extract at least one quotation related to at least one topic from content to be analyzed and perform quotation-based topic assignment, or other types of machine learning models including non-linear models and / or linear models, and may be composed of a combination thereof.

[0072] According to an embodiment of the present invention, a user computing device (1010) may perform a quotation-based topic assignment method using a large language model by using a local or / and external machine learning model (1020). Alternatively, the user computing device (1010) may perform a quotation-based topic assignment method using a large language model by using a machine learning model (1040) provided by a server.

[0073] In addition, according to another embodiment of the present invention, a server computing system (1030) communicating with a user computing device (1010) may provide a topic analysis result for content to be analyzed according to the present invention to the user computing device (1010) on an application or / and the web in accordance with a request from a user received through the user computing device (1010).

[0074] In addition, according to another embodiment of the present invention, by performing a quotation-based topic assignment method using a large language model of a user computing device (1010) and a server computing system (1030), a topic analysis result for the content to be analyzed can be provided to the user.

[0075] Additionally, according to various embodiments of the present invention, a user computing device (1010) and / or a server computing system (1030) can learn machine learning models (1020, 1040) performed in a quotation-based topic assignment method using a large language model through interaction with a training computing system (1050) that is communicatedly connected via a network (1070). In this case, the training computing system (1050) may be a computing system separate from the server computing system (1030). Alternatively, in some embodiments, the training computing system (1050) may be part of the server computing system (1030) or part of the user computing device (1010).

[0076] Meanwhile, the server computing system (1030) may include at least one processor (1031) and memory (1032). Here, the processor (1031) may be composed of at least one or a plurality of electrically connected processors among a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), a neural network processing unit (NPU), an application integrated circuit, an application semiconductor (ASIC), an arithmetic logic unit (ALU), a floating-point arithmetic unit (FPU), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, and / or other electrical units for performing functions. For example, at least one processor (1031) may include a circuit and a transistor configured to execute instructions from memory (1032).

[0077] The memory (1032) constituting the server computing system (1030) according to the present invention may include volatile memory, non-volatile memory, fixed media, removable media, magnetic media, optical media, semiconductor media, and / or other types of physically durable storage media. For example, the memory (1032) may include one or more non-transient / transient computer-readable storage media such as RAM, ROM, HDD (Hard Disk Drive), SSD (Solid State Disk), SSD (Silicon Disk Drive), EEPROM, EPROM, flash memory device, magnetic disk, etc., and combinations thereof, and may include web storage of a server that performs the storage function of memory over the internet. Additionally, the server computing system (1030) may further include a data storage (data store). For example, the data storage may be composed of at least one of a relational database, a NoSQL database, a data warehouse, and a local file system.

[0078] In the memory (1032) constituting the server computing system (1030) according to the present invention, data and instructions necessary for the at least one processor (1031) to perform the operation of an application for quotation-based topic assignment using a large language model may be stored.

[0079] In one embodiment, the server computing system (1030) may be composed of a single device or a plurality of computing devices, and these may be configured to operate according to a sequential or parallel computing architecture. Additionally, a distributed processing system may be configured with a plurality of networked devices.

[0080] Meanwhile, the training computing system (1050) may include at least one processor (1051) and memory (1052). The model trainer (1060) is a logical component that executes the training of at least one machine learning model (1020, 1040) and may be implemented in the form of hardware, firmware, or software. For example, the model trainer (1060) may be executed by the processor (1051) after loading training data (1061) stored in a storage device into memory (1052). For example, the model trainer (1060) may be configured to execute one or more operations (e.g., model training, model reconstruction, model validation, model testing) on ​​at least one machine learning model.

[0081] The machine learning model of the present invention may include at least one of a statistical model, an algorithm, a neural network (NN), a convolutional neural network (CNN), a generative neural network (GNN), a Word2Vec model, a Bag of Words model, a TF-IDF (document frequency-inverse document frequency) model, a GPT (Generative Pre-trained Transformer) model (or other autoregressive models), a PPO (Proximal Policy Optimization) model, a nearest neighbor model (e.g., a k-nearest neighbor model), a linear regression model, a K-means clustering model, a Q-learning model, a TD (Temporal Difference) model, a Deep Adversarial Network model, and all other types of models further described herein.

[0082] Specifically, the model trainer (1060) may execute operations to train a machine learning model, and said operations may include at least one of adding, removing, and modifying model parameters. At this time, the training of the machine learning model may be at least one of supervised learning, semi-supervised learning, and unsupervised learning. In one embodiment, the training of the machine learning model may include the step of repeatedly inputting training data (1061) based on epochs and repeatedly performing the machine learning model training process configured in this way. Here, an epoch may refer to a unit in which the entire set of training data (1061) undergoes forward and backpropagation processing once. In some implementations, different levels of training methods (e.g., supervised learning, semi-supervised learning, unsupervised learning) may be used for different epochs.

[0083] The training data (1061) of the present invention may include input data and / or data previously output from at least one machine learning model (e.g., recursive learning feedback).

[0084] At least one parameter of a machine learning model may include at least one of a seed value, a model node, a model layer, an algorithm, a function, connections between different machine learning models, connections between parameters, machine learning model constraints, and other digital components that influence the output of the machine learning model. In this case, model connections between different machine learning models may include or represent relationships between model parameters and / or models, which may be dependent or interdependent, hierarchical, and / or static or dynamic. The combinations and configurations of model parameters described herein may be too complex to be maintained or utilized by human cognitive abilities.

[0085] In the present invention, the machine learning parameters described according to the embodiments are not limited, and a single machine learning model may further include a plurality of model parameters.

[0086] Meanwhile, FIG. 2 illustrates an example of a block diagram of a computing device (1100) that may be included in a user computing device (1010), a server computing system (1030), and a training computing system (1050), as an embodiment of a computing system (10000) in which the present invention can be implemented.

[0087] As illustrated in FIG. 2, the computing device (1100) may include at least one application (e.g., Application 1 to Application N), and each of the at least one application may include a machine learning library and a model execution environment for performing a quotation-based topic assignment method using a machine learning-based large language model. The at least one application included in the computing device (1100) may communicate with the sensor, context manager, device state manager, or additional component(s) within the computing device (1100) via an Application Programming Interface (API). In one embodiment, the at least one application may interface with device components, such as receiving sensor data or state data or transmitting prediction results to an output device via a public or private API.

[0088] Meanwhile, FIG. 3 illustrates an example of a block diagram in another aspect of a computing device (1200), which is one of the components of a computing system (10000) that performs a quotation-based topic assignment method using a large language model according to an embodiment of the present invention.

[0089] A computing device (1200) according to the present invention may include at least one application (e.g., Application 1 to Application N), and at least one application may communicate with a central intelligence layer (1210). Each application may interact with a shared model within the central intelligence layer (1210) through an API (e.g., a common API).

[0090] The central intelligence layer (1210) includes one or more machine learning models and may share them among multiple applications or provide them independently to each. In one embodiment, the central intelligence layer (1210) may be integrated as part of an operating system or implemented as a separate logical layer.

[0091] Additionally, the central intelligence layer (1210) can communicate with the central device data layer (1220). The central device data layer (1220) can store the content to be analyzed and topic assignment information stored within the computing device (1200) and provide this as input data necessary for training (e.g., fine-tuning) at least one model to be trained. Each device component (e.g., sensor, state manager, etc.) can communicate with the central device data layer (1220) via a private API, etc.

[0092] The technology described herein may be composed of a single or multiple computing devices, and a machine learning model that performs a quotation-based topic assignment method using a large language model may be executed sequentially or in parallel on one component or multiple distributed components. Data storage, machine learning models, and applications may be distributed and operated locally or over a network, and these configurations can be flexibly applied to various system architectures.

[0093] Meanwhile, the present invention relates to a method and system for assigning at least one topic to content to be analyzed based on quotations using a large language model. Here, the content to be analyzed is data that can be provided or collected in text form and may refer to at least one content data collected (or received) from various sources (e.g., a web corpus, a document corpus, a database (DB) website, an API, a server linked to the topic assignment system (100), a central server, an external server, cloud storage, a user terminal, a large dataset, etc.). For example, the content to be analyzed may refer to content data including documents, news articles, reviews, blog posts, webpage content, chat logs, and product descriptions.

[0094] Meanwhile, the “Large Language Model” according to the present invention may refer to an artificial intelligence model capable of understanding and generating natural language by learning a vast amount of data. Specifically, the Large Language Model according to the present invention may refer to an artificial intelligence trained to identify at least one topic with respect to content to be analyzed and to extract a quotation corresponding to at least one topic.

[0095] In this case, according to the present invention, “Quotation” may refer to at least one set of words consisting of at least one word that is semantically consecutive in at least one sentence constituting the content to be analyzed. For example, a quotation according to the present invention may refer to at least one set of words (e.g., “It is a thing of beauty”) consisting of at least one word (e.g., “It,” “is,” “a,” “thing,” “of,” “beauty”) that is semantically connected in a specific content to be analyzed (e.g., “It is a thing of beauty and fast enough.”).

[0096] More specifically, in the present invention, a large language model can extract at least one quotation from the content to be analyzed and assign at least one topic on a quotation-by-quote basis.

[0097] In this case, according to the present invention, the “Topic” is configured to include at least one of a Domain and a Category, and may refer to subject information representing the core content of the text. Specifically, according to the present invention, the “Domain” may refer to subject information of a higher-level concept constituting the topic. For example, in the text “This laptop is lightweight.”, the domain may refer to “LAPTOP.”

[0098] In addition, the “category” according to the present invention may refer to subject information of a sub-concept constituting the topic. For example, in the text “This laptop is lightweight.”, the category may refer to “PORTABILITY.” To explain according to the example described above, the topic assigned to the quotation “This laptop is lightweight.” in the present invention may refer to “LAPTOP#PORTABILITY,” which is composed of at least one of the domain “LAPTOP” and the category “PORTABILITY.”

[0099] The present invention proposes a quotation-based topic assignment method and system using a large language model, which utilizes a large language model to assign topics in units of quotations rather than in units of entire documents within the content to be analyzed, thereby providing topic analysis results to a user terminal.

[0100] Meanwhile, the topic assignment system according to the present invention may include at least one artificial intelligence model. In this case, the at least one artificial intelligence model mentioned in the present invention may include various models that can be utilized depending on various situations or purposes. For example, the artificial intelligence model of the present invention may include at least one of a machine learning (ML) model, a deep learning model, a deep neural network (DNN), a language model (LM), a large language model (LLM), a super-large foundation model, a generative artificial intelligence (Generative AI) model, a transformer-based model, a supervised learning (SL) model, a reinforcement learning (RL) model, and a special purpose model (e.g., a time series forecasting model (e.g., ARIMA model, SARIMA model, etc.), a time series foundation model, a graph neural network (GNN), a multimodal model, a natural language processing (NLP) model, a computer vision model, a speech recognition / synthesis model, a recommendation system model, etc.).

[0101] Furthermore, the artificial intelligence model used in the topic assignment system according to the present invention may be implemented with a single or multiple components. More specifically, the artificial intelligence model used in the present invention may be implemented with a single component (1) or with multiple components (2, 3, 4 or more, etc.).

[0102] In this regard, when multiple artificial intelligence models are implemented in the present invention, the multiple artificial intelligence models may be implemented as models of the same type. In this case, the multiple artificial intelligence models may be implemented to perform different functions (or roles) in the present invention.

[0103] In the foregoing, a quotation-based topic assignment method using a large language model according to the present invention has been generally described, and this can be implemented by a quotation-based topic assignment system using a large language model described below. Below, with reference to FIGS. 4 and FIGS. 5, a quotation-based topic assignment system using a large language model according to the present invention will be described in detail. FIGS. 4 is a conceptual diagram for explaining a quotation-based topic assignment system using a large language model according to the present invention, and FIGS. 5 is a conceptual diagram for explaining concepts related to quotation-based topic assignment according to the present invention.

[0104] Meanwhile, as illustrated in FIG. 4, a quotation-based topic assignment system using a large language model according to the present invention (hereinafter referred to as the “topic assignment system,” 100) may include at least one of a communication unit (110), a storage unit (120), and a control unit (130). However, the components of the topic assignment system (100) according to the present invention are not limited thereto and may further include various hardware components that perform the same or similar roles as described in the description of the present specification.

[0105] Although not illustrated, the topic assignment system (100) according to the present invention may include one or more processors, and such processors may include one or more general-purpose processors and / or one or more special-purpose processors (e.g., digital signal processor, tensor processing unit (TPU), graphics processing unit (GPU), neural network processing unit (NPU), application integrated circuit, application semiconductor (ASIC), field programmable gate array (FPGA), quantum processing unit (or quantum processor, QPU), etc.). One or more processors may be configured to execute instructions, computer-readable instructions, and / or other instructions described herein that are stored (or included) in the storage unit (120).

[0106] In the quotation-based topic assignment method and system using a large language model according to the present invention, a memory and at least one processor can cooperate to perform the data processing described below. The processor can perform a series of operations and data processing using data and information stored in the memory. At this time, the memory may be a component of the storage unit (120).

[0107] In addition, the topic assignment system (100) according to the present invention can perform data processing and computation processes using quantum gates, quantum entanglement, and quantum superposition states, taking into consideration implementation in a quantum computer environment. For example, the present invention can perform parallel computations based on qubits, and such quantum computations can operate complementarily with existing classical computers.

[0108] Such quantum computers may include parallel computation using qubits and high-speed data processing devices utilizing quantum entanglement, and hardware-based computational optimization using FPGAs and ASICs is possible. In addition, quantum computers may utilize quantum processors capable of qubit-based parallel computation, and data processing efficiency can be improved through a hybrid structure with existing classical computers.

[0109] Meanwhile, the topic assignment system (100) according to the present invention may exist inside a server (hereinafter referred to as a server) built to perform a specific purpose (e.g., topic assignment), or it may exist as a separate system from said server. When the topic assignment system (100) exists inside the server, it may provide various services related to the present invention (e.g., providing topic analysis results and topic analysis content, etc.) through at least one component located inside the server or configuration modules that perform functions similar to said components. In this case, the application may provide various services related to the present invention on an electronic device on which the application is installed through communication with the server.

[0110] Meanwhile, the communication unit (110) according to the present invention may be connected to a user terminal (10), an LLM server (140), a database (200), a central server, a device, and at least one network via a wireless or wired network, and may be configured to receive or transmit overall data and information necessary for the operation of the topic allocation system (100) according to the present invention.

[0111] The communication unit (110) can receive content to be analyzed from at least one source (or source server). As previously described, the communication unit (110) can receive at least one content data collected (or received) from various sources (e.g., a web corpus, a document corpus, a database (DB) website, an API, a server linked to the topic assignment system (100), a central server, an external server, cloud storage, a user terminal, a large dataset, etc.) as content to be analyzed.

[0112] Meanwhile, the communication unit (110) can receive a user query from the user terminal (10). Here, “receiving a user query” may mean receiving an input signal corresponding to the user query based on the user query being input by the user through the input unit configuration provided in the user terminal (10).

[0113] Here, the user terminal (10) may include at least one of a mobile phone, a smartphone, a notebook computer, a laptop computer, a slate PC, a tablet PC, an ultrabook, a desktop computer, a digital broadcasting terminal, a PDA (personal digital assistants), a PMP (portable multimedia player), a navigation device, and a wearable device (e.g., a smartwatch, a smart glass, a head-mounted display).

[0114] In addition, the input section of the user terminal (10) in the present invention does not necessarily refer to a hardware means, but can be understood as a channel for receiving input from a user. Such an input section of the user terminal (10) may also be referred to as a user interface module. The input section of the user terminal (10) may include a touch screen, a computer mouse, a keyboard, a keypad, a touchpad, a trackball, a joystick, a voice recognition module, or other similar devices. However, the present invention does not limit the type of input section of the user terminal (10). Here, user input may include documents, text, images (or videos), voice, etc. In this case, the topic assignment system (100) may further include a module for converting voice into text.

[0115] Meanwhile, the communication unit (110) may include at least one communication module capable of wireless communication and wired communication between the topic assignment system (100) and the communication target. Additionally, the communication unit (110) may include a communication module that connects the topic assignment system (100) to at least one network.

[0116] Meanwhile, the communication unit (110) can support various communication methods depending on the communication standard of the device being communicated. For example, the communication unit (110) can be configured to communicate with a communication target using at least one of the following technologies: WLAN (Wireless LAN), Wi-Fi (Wireless-Fidelity), Wi-Fi (Wireless Fidelity) Direct, DLNA (Digital Living Network Alliance), WiBro (Wireless Broadband), WiMAX (World Interoperability for Microwave Access), HSDPA (High Speed ​​Downlink Packet Access), HSUPA (High Speed ​​Uplink Packet Access), LTE (Long Term Evolution), LTE-A (Long Term Evolution-Advanced), 5G (5th Generation Mobile Telecommunication), Bluetooth (Bluetooth™ RFID (Radio Frequency Identification), Infrared Data Association (IrDA), UWB (Ultra-Wideband), ZigBee, NFC (Near Field Communication), and Wireless USB (Wireless Universal Serial Bus).

[0117] Next, the storage unit (120, or memory) serves to store various data related to the present invention, and the storage unit (120) may be provided in the topic allocation system (100) itself, or alternatively, at least part of the storage unit (120) may mean a database (DB) or memory. The storage unit (120) may include one or more non-transient computer-readable storage media that can be read and / or accessed by at least one processor. One or more computer-readable storage media may include volatile and / or non-volatile storage components such as optical, magnetic, organic, or other memory or disk storage devices. In some examples, the storage unit (120) may be implemented using a single physical device (e.g., one optical, magnetic, organic, or other memory or disk storage device), whereas in other examples, the storage unit (120) may be implemented using multiple physical devices.

[0118] The storage unit (120) may include computer-readable instructions and additional data. The storage unit (120) may include a storage necessary to perform at least some of the methods, scenarios, and techniques described in this specification and / or at least some of the functions of the device and network. Furthermore, at least some of the storage unit (120) may be a cloud storage or a cloud server. That is, the storage unit (120) is sufficient as a space where information necessary for the operation of the topic assignment system (100) according to the present invention is stored, and it can be understood that there are no restrictions on the physical space. Accordingly, the storage unit (120) and the database (200) may be used interchangeably without being separately distinguished below.

[0119] Furthermore, at least a portion of the storage unit (120) may be a cloud storage or a cloud server. That is, the storage unit (120) is sufficient as a space where information necessary for the operation of the topic allocation system (100) according to the present invention is stored, and it can be understood that there are no restrictions on the physical space.

[0120] The storage unit (120) may store at least one content to be analyzed collected (or received) from various sources (e.g., a web corpus, a document corpus, a database (DB) website, an API, a server linked to the topic assignment system (100), a central server, an external server, cloud storage, a user terminal (10), a large dataset, etc.). For example, the storage unit (120) may store content to be analyzed including documents, news articles, reviews, blog posts, webpage content, chat logs, and product descriptions received from various source servers.

[0121] The content to be analyzed stored in the storage unit (120) according to the present invention is not limited to the examples described above and may include all types of data that are expressed in text form or convertible into text.

[0122] Meanwhile, the storage unit (120) may store user information related to multiple users registered in the topic assignment system (100). For example, the storage unit (120) may store user information including at least one of each of the following: i) name, ii) age (or age group), iii) gender, iv) interests, v) occupation or job information, vi) location information (e.g., residential area, country), vii) service usage history, viii) past query information, ix) search history, x) past topic analysis result history, xi) user settings (preferences) or personalization options, xii) login account or identification information, and xiii) device information (e.g., type of device used).

[0123] The user information stored in the storage unit (120) according to the present invention is not limited to the examples described above, and may further include various forms of information that can be utilized to identify or analyze at least one of the user's identity, environment, preference, and service usage behavior in the process of providing topic analysis content related to the content to be analyzed.

[0124] Meanwhile, data and commands necessary for the operation of the topic assignment system (100) according to the present invention may be stored in the storage unit (120). Specifically, commands for the operation of the prompt generation unit (131) may be stored in the storage unit (120). For example, the prompt generation unit (131) according to the present invention may refer to a module capable of generating a prompt that requests to generate a topic analysis result, by extracting at least one of at least one topic candidate and at least one quotation related to the content to be analyzed.

[0125] For example, the prompt generation unit (131) may include at least one of a large language model based on T5 (Text-to-Text Transfer Transformer), BART (Bidirectional and Auto-Regressive Transformer), GPT (Generative Pre-trained Transformer), or LLaMA (Language Model for Many Applications), a rule-based template matching algorithm, a conditional prompting technique, a contextual embedding selection module, or a few-shot prompt generator.

[0126] At this time, the model or algorithm included in the prompt generation unit (131) according to the present invention is not limited to the examples described above, and may further include artificial intelligence models and algorithms that perform the same function.

[0127] Furthermore, the storage unit (120) can store a computer program including computer program instructions. Furthermore, the storage unit (120) can store a computer program including computer program instructions that control the operation of the topic allocation system (100) or control the operation of the control unit (130) when loaded into the processor of the topic allocation system (100).

[0128] Next, the control unit (130) can perform the role of controlling the overall operation of the topic assignment system (100) related to the present invention. Specifically, the control unit (130) may include a prompt generation unit (131). In this case, although the present invention describes the control unit (130) as including the prompt generation unit (131), it is not limited thereto. If the prompt generation unit (131) exists on an external server, the control unit (130) may call or control the function of the prompt generation unit (131) by linking with the external server. For example, if the prompt generation unit (131) according to the present invention exists on an LLM server (140), the control unit (130) may call or control the function of the prompt generation unit (131) by linking with the LLM server (140).

[0129] At this time, although the present invention describes the LLM server (140) as existing separately from the topic allocation system (100), it is not limited thereto, and the topic allocation system (100) may be configured to include the LLM server (140). That is, the topic allocation system (100) and the LLM server (140) according to the present invention may exist separately, or the LLM server (140) may be included in the topic allocation system (100). For convenience of explanation, the LLM server (140) and at least one large language model are used interchangeably below, and the use of the large language model by the control unit (130) can be understood as using at least one large language model included in the LLM server (140).

[0130] The control unit (130) can process signals, data, information, etc. that are input or output through the components of the topic assignment system (100) described above, or perform a series of data processing to provide or process appropriate information and functions to the user. The control unit (130) can be physically implemented by the processor described above.

[0131] As illustrated in FIG. 5(a), the control unit (130) can generate a topic analysis result for the content to be analyzed according to a quotation-based topic allocation method (QBTA) rather than a document-based topic allocation method (DBTA) for allocating topics in sentence units in the content to be analyzed.

[0132] Specifically, the control unit (130) does not assign topics based on the proportion of topics in the entire content unit to be analyzed, as in the DBTA method, but rather uses a large language model to identify at least one semantically cohesive quotation within a sentence and assigns a topic that directly corresponds to at least one quotation, thereby allowing it to select and display only text sections that are substantially associated with at least one topic.

[0133] To this end, the control unit (130) can specify an analysis target content consisting of at least one sentence and at least one topic related to the at least one sentence constituting the analysis target content.

[0134] For example, the control unit (130) can use a large language model (141) to generate at least one topic candidate based on the content to be analyzed, and based on the at least one topic candidate, identify at least one topic related to at least one sentence.

[0135] As another example, the control unit (130) can identify at least one topic related to at least one sentence based on topic candidate information stored in advance in the database (200). In this case, the topic candidate information may refer to topic information stored by classifying and defining according to at least one of a domain and a category. For example, the database (200) may store at least one topic candidate information that is a combination of at least one of domain information representing a specific field (e.g., LAPTOP, RESTAURANT, NEWS, etc.) and category information representing a sub-topic within the domain (e.g., DESIGN_FEATURES, PERFORMANCE, PRICE, etc.).

[0136] As illustrated in FIG. 5(b), the control unit (130) can generate at least one topic candidate (e.g., capacity, skin improvement, price) from at least one sentence (e.g., “It has a generous capacity and is good for the skin, and I like it because the price is reasonable”) that constitutes the content to be analyzed, and can specify at least one topic based on the at least one topic candidate.

[0137] Furthermore, the control unit (130) can use a large language model to extract at least one quotation corresponding to each of at least one topic from the content to be analyzed. For example, the control unit (130) can extract at least one quotation (e.g., “the capacity is generous and it is good for the skin, the price is reasonable and I like it”) corresponding to each of at least one topic candidate (e.g., capacity, skin improvement, price) from at least one sentence (e.g., “the capacity is generous and it is good for the skin, the price is reasonable and I like it”) constituting the content to be analyzed.

[0138] Specifically, the control unit (130) can generate a prompt requesting the extraction of at least one quotation based on the content to be analyzed and at least one topic candidate. Furthermore, the generated prompt can be processed as input to a large language model (141) to extract at least one quotation from the content to be analyzed from the large language model.

[0139] Furthermore, the control unit (130) can provide a topic analysis result for the content to be analyzed using topics and quotations. Here, the topic analysis result may refer to data containing at least one topic assignment information in which a specific topic is mapped to at least one specific quotation. Specifically, the control unit (130) can generate a topic analysis result configured such that, based on at least one topic assignment information, at least one quotation included in the content to be analyzed is classified and displayed according to at least one topic, using a large language model.

[0140] Furthermore, the control unit (130) may provide the generated topic analysis result and topic analysis content related to the topic analysis result to the user terminal (10). Here, the topic analysis content may refer to content including at least one of the topic analysis result corresponding to the user query and response information to the user query.

[0141] Specifically, the control unit (130) can receive a user query related to a document to be analyzed from a user terminal (10). In the present invention, receiving a user query from a user terminal (10) can be interpreted as having the meaning of “receiving a user query through at least one page output to the user terminal (10)” or “receiving a user query from a user terminal (10) on which at least one page has been output.”

[0142] Furthermore, the control unit (130) can use a large language model to generate topic analysis content based on user queries and content to be analyzed, and provide the generated topic analysis content to the user terminal (10).

[0143] Meanwhile, the topic assignment system (100) may include one or more processors, and such processors may include one or more general-purpose processors and / or one or more special-purpose processors (e.g., digital signal processors, tensor processing units (TPUs), graphics processing units (GPUs), neural network processing units (NPUs), application integrated circuits, application semiconductors (ASICs), etc.). One or more processors may be configured to execute instructions, computer-readable instructions, and / or other instructions described herein that are stored (or included) in the storage unit (120). Such a topic assignment system (100) may perform data processing described below in cooperation with memory and at least one processor. The processor may be electrically connected to memory and may perform a series of operations and data processing using data and information stored in memory. Here, “memory” may be a component of the storage unit (120), and “processor” may be used interchangeably with the control unit (130).

[0144] In the above description, the topic assignment system (100) of the present invention has been described, and it can implement a quotation-based topic assignment method using a large language model described below.

[0145] Hereinafter, with reference to FIG. 6 together with FIG. 7, FIG. 8a and FIG. 8b, FIG. 9, FIG. 10, and FIG. 11a to FIG. 11d, a quotation-based topic assignment method using a large language model according to the present invention will be described in more detail. FIG. 6 is a flowchart for explaining a quotation-based topic assignment method using a large language model according to the present invention, FIG. 7, FIG. 8a, FIG. 8b, FIG. 9 and FIG. 10 are conceptual diagrams for specifically explaining a quotation-based topic assignment method using a large language model according to the present invention, and FIG. 11a to FIG. 11d are conceptual diagrams for explaining various embodiments using a quotation-based topic assignment method using a large language model according to the present invention.

[0146] Meanwhile, in the present invention, a process of specifying the content to be analyzed, consisting of at least one sentence, may be performed (S610, see FIG. 6).

[0147] The control unit (130) can collect at least one content data collected (or received) from various sources (e.g., a web corpus, a document corpus, a database (DB) website, an API, a server linked to the topic assignment system (100), a central server, an external server, cloud storage, a user terminal, a large dataset, etc.) as content to be analyzed.

[0148] Specifically, the control unit (130) receives at least one content data including content to be analyzed from a user terminal (10), and can identify the content to be analyzed in relation to a user query entered into the user terminal.

[0149] For example, the control unit (130) may provide at least one page to the user terminal (10) that includes various functions for collecting content to be analyzed, which consists of at least one sentence. The control unit (130) may receive at least one content data based on (or derived from) user input received from the user terminal (10), and may identify at least one content to be analyzed from the at least one content data.

[0150] As another example, the control unit (130) may collect at least one content data from an external source server and identify at least one content to be analyzed from the at least one content data. For example, the control unit (130) may collect at least one content data from one of a website server, a portal site server, a blog / cafe platform server, a community server, a news article providing server, a social media (SNS) server, a video platform server, a shopping mall server, a review platform server, an internal document management server (DMS), a cloud storage server, a database (DB) server, and an Open API server. Specifically, the control unit (130) may collect at least one content data using at least one method among web crawling and web scraping, and identify at least one content to be analyzed from the at least one content data.

[0151] At this time, the method for specifying the content to be analyzed in the present invention is not limited to the examples described above, and can be implemented in a wide variety of ways so that various modifications are possible depending on at least one of the type of content to be analyzed, the collection path, and the user query.

[0152] That is, in the present invention, the control unit (130) may specify the content to be analyzed based on at least one content data directly input from a user terminal, may specify it based on at least one content data collected from an external source server, or may specify it based on at least one content data retrieved from a previously stored database. Furthermore, it can be understood that the control unit (130) may specify the content to be analyzed consisting of at least one sentence by using all methods, whether using these methods alone or in combination with each other.

[0153] As illustrated in FIG. 7, the control unit (130) can identify the content to be analyzed (700) from the at least one content data. For example, the control unit (130) can identify the content to be analyzed (700) from at least one of at least one document (701), at least one web page (702), at least one link content (703), and at least one application (704).

[0154] Specifically, the control unit (130) may identify all types of data that are expressed in text form or convertible into text in at least one of at least one document (701), at least one web page (702), at least one link content (703), and at least one application (704) as the content to be analyzed (700). At this time, the content to be analyzed (700) according to the present invention may consist of at least one sentence.

[0155] For example, the control unit (130) may provide at least one page to the user terminal (10) that includes various functions for collecting content to be analyzed, which consists of at least one sentence. The control unit (130) may receive at least one content to be analyzed based on (or derived from) user input received from the user terminal (10).

[0156] The user input according to the present invention may include at least one of multimodal input data comprising text, images, videos, voice, and combinations thereof, which include content to be analyzed. Specifically, the control unit (130) receives at least one content data according to the user input received from the user terminal (10), and can identify at least one content to be analyzed from the at least one content data.

[0157] Next, in the present invention, a process may be carried out in which at least one topic related to at least one sentence constituting the content to be analyzed is identified (S620, see FIG. 6).

[0158] As previously explained, the “Topic” according to the present invention is configured to include at least one of a Domain and a Category, and may refer to subject information representing the core content of the content to be analyzed. Specifically, the “Domain” according to the present invention may refer to subject information of a higher-level concept constituting the said Topic. Additionally, the “Category” according to the present invention may refer to subject information of a lower-level concept constituting the said Topic.

[0159] The control unit (130) can identify topic information representing the core content of the content to be analyzed in relation to at least one sentence constituting the content to be analyzed. For example, the control unit (130) can generate at least one topic candidate from the content to be analyzed using a large language model (141). As illustrated in FIG. 8a, the control unit (130) can generate a first prompt (800) requesting the generation of at least one topic candidate consisting of at least one of a domain and a category based on the content to be analyzed (700) using a prompt generation unit (131).

[0160] As illustrated in FIG. 8b, the control unit (130) can process the first prompt (800) as input to the large language model (141) to obtain at least one topic candidate (710) from the large language model (141). In another example, the control unit (130) can identify at least one topic related to at least one sentence based on topic candidate information stored in advance in the database (200). Specifically, the large language model (141) can obtain at least one topic candidate (710) related to at least one sentence constituting the content to be analyzed by referring to at least one topic candidate information that is a combination of at least one of domain information (720) representing a specific field (e.g., LAPTOP, RESTAURANT, NEWS, etc.) stored in advance in the database (200) and category information (730) representing a sub-topic within the domain (e.g., DESIGN_FEATURES, PERFORMANCE, PRICE, etc.).

[0161] Furthermore, the control unit (130) can identify at least one topic related to the at least one sentence based on at least one topic candidate (710) obtained. In the present invention, the method of identifying at least one topic related to the at least one sentence based on at least one topic candidate can be very diverse. For example, the control unit (130) can identify all of the at least one topic candidate obtained as at least one topic.

[0162] As another example, the control unit (130) may use a large language model (141) to select at least some of at least one topic candidate and specify the selected at least some of the topic candidates as at least one topic associated with at least one sentence. Specifically, the control unit (130) may use a prompt generation unit (131) to generate a first prompt configured to generate at least some of the topic candidates among the generated at least one topic candidate. At this time, the pre-set number is not limited to a specific number and can be set in various ways according to at least one of the system settings, user settings, and the content of the first prompt, and can be changed dynamically.

[0163] The control unit (130) can obtain at least some topic candidates from the large language model (141) by processing a first prompt configured to generate at least some topic candidates as a preset number among at least one generated topic candidate as input to the large language model (141). Furthermore, the control unit (130) can identify at least some topic candidates corresponding to the preset number as at least one topic related to at least one sentence constituting the analysis content.

[0164] In the present invention, all of at least one topic candidate may be designated as at least one topic, or only some of the topic candidates may be selected and designated as at least one topic. Accordingly, in the present invention, 'at least one topic' and 'at least one topic candidate' may be used interchangeably depending on the technical context.

[0165] As such, the method for identifying at least one topic related to at least one sentence constituting the content to be analyzed in the present invention can be very diverse. The present invention is not limited to artificial intelligence models, topic candidate calculation methods, predefined domain and category structures, etc., but can be understood to include all methods capable of deriving at least one topic related to at least one sentence constituting the content to be analyzed by utilizing at least one of a generation method using a large language model, a query method using topic candidate information stored in a database, or an inference method through linkage with an external knowledge base.

[0166] Next, in the present invention, a process of extracting at least one quotation corresponding to each of at least one topic from the content to be analyzed using a large language model may be performed (S630, see FIG. 6).

[0167] A “quotation” according to the present invention may mean a set of at least one word composed of at least one word that is semantically consecutive in at least one sentence. For example, a quotation according to the present invention may mean a set of at least one word (e.g., “It is a thing of beauty”) composed of at least one word (e.g., “It,” “is,” “a,” “thing,” “of,” “beauty”) that is semantically connected in at least one sentence (e.g., “It is a thing of beauty and fast enough.”) constituting a specific content subject to analysis.

[0168] The control unit (130) can extract at least one quotation according to at least one topic related to at least one sentence constituting the content to be analyzed. As illustrated in FIG. 9, the control unit (130) can use the prompt generation unit (131) to generate a second prompt (900) requesting the extraction of at least one quotation based on the content to be analyzed (700) and at least one topic candidate (or topic, 710). Specifically, the prompt generation unit (131) can configure the second prompt (900) so that the large language model (141) extracts at least one set of words consisting of at least one word that is semantically consecutive in at least one sentence as at least one quotation.

[0169] For example, the prompt generation unit (131) may generate a second prompt (900) including content (901) that assigns a role for topic assignment to a large language model (141) in order to extract a quote corresponding to at least one topic related to at least one sentence constituting the content to be analyzed.

[0170] Additionally, the control unit (130) may generate a second prompt (900) including content (902) commanding the extraction of quotations. Additionally, the control unit (130) may generate a second prompt (900) including content (903) regarding conditions for the extraction of quotations. Additionally, the control unit (130) may generate a second prompt (900) including content (904, 905) requesting the extraction of quotations based on at least one of the analysis target content, the domain and category of at least one specified topic.

[0171] At this time, the content included in the second prompt (900) in the present invention is not limited to the examples described above, and may be modified to include various additional conditions for controlling the response method of a large language model, such as assigning priority to a specific topic. In addition, the configuration of the second prompt (900) may be optionally added, deleted, or replaced according to a user request.

[0172] The control unit (130) processes the second prompt (900) as input to the large language model (141) to extract at least one quotation from the content to be analyzed from the large language model (141). Furthermore, the control unit (130) can use the large language model (141) to assign a specific topic among at least one topic to each of the at least one quotation.

[0173] Specifically, the control unit (130) can generate a second prompt (900) configured to map a specific topic among at least one topic to at least one specific quote corresponding to the specific topic using a prompt generation unit (131). At this time, mapping a specific topic among at least one topic to at least one specific quote corresponding to the specific topic may mean defining the correspondence relationship between the topic and the quote in a structured form so that the specific quote can be used as evidence to explain or support the specific topic.

[0174] That is, the control unit (130) can generate a second prompt requesting the large language model (141) to map a specific topic among at least one topic and at least one specific quote corresponding to the specific topic.

[0175] Next, in the present invention, a process of providing a topic analysis result for the content to be analyzed using the topic and the quotation may be carried out (S640, see FIG. 6).

[0176] As illustrated in FIG. 10, the control unit (130) can process the second prompt (900) as input to the large language model (141). Furthermore, the large language model (141) can output at least one topic assignment information (921 to 923) in which a specific topic is mapped to at least one specific quotation. Here, the topic assignment information (921 to 923) may refer to information in which the large language model (141) maps a specific topic among at least one topic to at least one specific quotation corresponding to the specific topic, thereby expressing the correspondence relationship between the topic and the quotation in the form of structured data. That is, it may include information indicating which quotation is used as evidence for the specific topic.

[0177] The control unit (130) can obtain a topic analysis result (910) for the content to be analyzed (700) from a large language model (141). At this time, the topic analysis result (910) can be configured so that at least one quotation (911, 912, 913) included in the content to be analyzed (700) is classified and displayed according to at least one topic based on at least one topic assignment information (921 to 923).

[0178] For example, the quote “It is a thing of beauty” included in the content (700) to be analyzed can be mapped to the topic “LAPTOP#DESIGN_FEATURES”, which consists of the domain “LAPTOP” and the category “DESIGN_FEATURES”. Additionally, the quote “fast enough” can be classified to correspond to the topic “LAPTOP#OPERATION_PERFORMANCE”. Furthermore, the quote “The battery will get you from LA to NY no problem” can be displayed to correspond to the topic “BATTERY#OPERATION_PERFORMANCE”, which consists of the domain “BATTERY” and the category “OPERATION_PERFORMANCE”.

[0179] In this way, the control unit (130) can obtain a displayed topic analysis result by clearly distinguishing which of at least one topics a quote extracted from a large language model (141) corresponds to.

[0180] At this time, the topic analysis result (910) may be configured to include quotation reference information to identify where a specific quotation was extracted from the content to be analyzed. Here, the quotation reference information may refer to reference information for identifying the section in the content to be analyzed where the quotation exists, including at least one of the start position and end position of the quotation within the content to be analyzed, the sentence number, and the offset within the sentence.

[0181] Specifically, the control unit (130) can configure the second prompt so that the topic analysis result (910) includes quotation reference information related to the location of at least one quotation included in at least one topic assignment information in the content to be analyzed. Furthermore, the control unit (130) can process the second prompt as input to a large language model and obtain the topic analysis result (910) including quotation reference information from the large language model (141).

[0182] Furthermore, the control unit (130) can provide the generated topic analysis result (910) to the user terminal (10). Specifically, the control unit (130) can provide the topic analysis result (910) to the user terminal (10) through at least one page associated with the topic allocation system (100).

[0183] In the present invention, the method by which the control unit (130) provides the topic analysis result (910) to the user terminal (10) can be very diverse. For example, the control unit (130) can provide the topic analysis result (910) to the user terminal (10) by rendering the generated response on a web page through a web-based interface.

[0184] As another example, the control unit (130) may be linked with a mobile application (App) or a dedicated client program to provide the topic analysis result (910) for the content to be analyzed in at least one of text, code blocks, and graphic visualization (UI Component).

[0185] As another example, the control unit (130) may provide the topic analysis result (910) for the content to be analyzed in the form of a natural language conversation with the user through a conversational chatbot interface, or provide the topic analysis result (910) for the content to be analyzed to the user terminal (10) through at least one of email, message and notification.

[0186] Additionally, the control unit (130) may convert the topic analysis result (910) for the content to be analyzed into a downloadable file format (e.g., .pdf, .docx, .py, .js, etc.) and provide it to the user terminal (10), and may also upload the topic analysis result (910) for the content to be analyzed to a project document or code repository by linking with a cloud storage and a collaboration platform (e.g., GitHub, Notion, Google Drive, etc.).

[0187] That is, the control unit (130) according to the present invention can provide the topic analysis result (910) to the user in various ways, such as web, app, interactive interface, or file and link-based, depending on the format of the content to be analyzed, the type of user terminal, and the user access environment.

[0188] The above describes a quotation-based topic assignment method using a large language model according to the present invention. In this regard, an embodiment using the quotation-based topic assignment method using a large language model will be described below with reference to FIGS. 11a to 11d. FIGS. 11a to 11d are conceptual diagrams for explaining various embodiments using the quotation-based topic assignment method using a large language model according to the present invention.

[0189] The control unit (130) can generate topic analysis content from the content to be analyzed according to a user query received from a user terminal.

[0190] As illustrated in FIG. 11a, the control unit (130) can receive a user query (932) requesting at least one operation using the topic analysis result for the content to be analyzed from the user terminal.

[0191] For example, the control unit (130) can retrieve user information (931) stored in the database (200) using an authentication token or identifier (ID) of a user account logged into the user terminal (10). As another example, when the control unit (130) is linked with a web-based service, it may collect user profile information by performing API communication with an external source server (e.g., SNS account server, integrated login server, cloud account server, etc.). Furthermore, the control unit (130) can infer user interest or preferred topic information by analyzing service usage patterns such as UI event logs, click patterns, search history, and past query records generated from the user terminal (10). At this time, it can be understood that the method of collecting user information in the present invention is not limited to the examples described above, and can be expanded in various ways by linking with at least one of the database (200), the user terminal (10), and an external server.

[0192] Furthermore, the control unit (130) can generate topic analysis content (940) corresponding to a user query (932) using a large language model (141). At this time, the topic analysis content (940) may be configured to include at least one of a topic analysis result corresponding to the user query (932) and response information for the user query (932).

[0193] It can be understood that the format of the topic analysis content (940) according to the present invention is not limited to a specific format and can be configured in various formats according to user query, such as text information, structured information in a table format, image-based visualization information, summary results, keypoint lists, and topic information in the form of tags.

[0194] Specifically, the control unit (130) can generate a prompt (or a third prompt) requesting the generation of topic analysis content (940) based on a user query (932) and content to be analyzed (700). At this time, the control unit (130) can use a prompt generation unit (131) to collect user information (931) corresponding to a user account logged into the user terminal (10), and generate a prompt (or a third prompt) requesting a large language model (141) to generate topic analysis content (940) based on the user information (931) and the user query (932).

[0195] Specifically, the control unit (130) can generate a prompt (or a third prompt) requesting the generation of topic analysis content (940) corresponding to the user query (932) by referring to collected user information (931) based on at least one of the user query (932) and the content to be analyzed (700) received from the user terminal (10).

[0196] Furthermore, the control unit (130) can input the generated prompt (or third prompt) into the large language model (141) to obtain topic analysis content (940) related to the user query (732) from the large language model (141). Specifically, the control unit (130) can process the prompt (or third prompt) requesting the generation of topic analysis content (940) as input to the large language model (141). Based on the prompt, the control unit (130) can obtain the topic analysis content (940) generated by the large language model (141) from the large language model (141).

[0197] As illustrated in FIG. 11b, the control unit (130) may provide the generated topic analysis content (940) to the user terminal (10) as a response to a user query (732). As previously described, the topic analysis content (940) according to the present invention may be configured to include at least one of a topic analysis result (910) corresponding to the user query (732) and a response information (941) including summary information related to the user query (932) extracted from the topic analysis result (910). At this time, the response information (941) according to the present invention is a response result based on a large language model for the user query (932) received from the user terminal (10), and may include various content generated according to the purpose of the user query.

[0198] For example, the control unit (130) can generate topic analysis content (940) containing specific response information (e.g., “Generate personalized topic analysis results and customized marketing copy in JSON based on this text”, tone and manner: emotional and practical tone) in response to receiving a specific user query from a user terminal (10) (e.g., “Generate personalized topic analysis results and customized marketing copy in JSON based on this text”, tone and manner: emotional and practical tone) in conjunction with a large language model (141).

[0199] As another example, as illustrated in FIG. 11c, the control unit (130) can use the prompt generation unit (131) to obtain a topic analysis result from the content to be analyzed (or the first content to be analyzed, 700a) and generate a prompt (or the third prompt) containing conditions included in the user query (932a) (e.g., “Show only low-rated reviews,” “Visualize with a pie chart”) and process it as input to the large language model (141).

[0200] Furthermore, the large language model (141) can select at least one quote (952, 953, 954, etc.) according to condition information included in the prompt (or third prompt) and filter only the quotes that correspond to the user query (932a). Specifically, the control unit (130) can generate response information (950) configured based on at least one quote mapped to at least one topic (951 to 954) according to condition information from the large language model (141).

[0201] As an example, the control unit (130) may generate response information (950) corresponding to a visual object (e.g., a pie chart) that represents the weight of each of at least one topic for the first analysis target content, based on the topic distribution of selected quotes (962, 963, 964, etc.). At this time, the response information (950) may be configured to include a list of representative phrases by topic (951) and topics mapped to each quote (e.g., delivery, noise, packaging condition, filter management, etc., 951 to 954).

[0202] It can be understood that the form of the response information (941) according to the present invention is not limited to a specific type and can be generated in various ways, such as text, table, JSON structure, visual representation (e.g., image description data), document form, etc., depending on the type and purpose of the user query.

[0203] Furthermore, the control unit (130) can provide the generated topic analysis content (940) to the user terminal (10). At this time, the method of providing the generated topic analysis content (940) to the user terminal (10) can be very diverse, and the present invention does not specifically limit the method of providing the generated topic analysis content (940) to the user terminal (10). For example, the control unit (130) can directly render and provide text-based analysis results through the screen (UI) of the user terminal (10). In addition, the control unit (130) can convert the generated topic analysis content (940) into a structured data format such as JSON, XML, or HTML and transmit it to the user terminal (10) in the form of an API response.

[0204] As another example, the control unit (130) may generate and provide a screen composed of various visual UI elements, such as graphs, charts, highlight displays, and card-shaped information blocks, so that they can be visually displayed in an application or web browser of the user terminal (10). Additionally, when using a voice-based interface, the control unit (130) may convert the topic analysis content (940) into voice synthesis data and provide it to the user terminal (10) in the form of voice guidance.

[0205] Furthermore, the control unit (130) can provide the generated topic analysis content (940) in the form of a push notification, message data, or email by utilizing the notification function, message function, email transmission function, etc. of the user terminal (10). As such, it can be understood that the method of providing the generated topic analysis content (940) to the user terminal (10) in the present invention can be implemented in various ways depending on the service type, user environment, device characteristics, etc.

[0206] In this way, the control unit (130) can generate user-customized topic analysis content (940) that takes into account user information (931) according to user query (932) and analysis target content (700), and provide it to the user terminal (10).

[0207] Meanwhile, the control unit (130) can determine a main topic and a sub-topic among at least one topic related to the content to be analyzed. At this time, the main topic may refer to topic information containing core content in the content to be analyzed, and the sub-topic may refer to topic information containing additional content in the content to be analyzed.

[0208] In the present invention, the method of determining the main topic and subtopic among at least one topic can be very diverse. For example, the control unit (130) can determine the main topic based on the number of at least one citation mapped to each of at least one topic. Furthermore, the control unit (130) can determine the remaining topics excluding the main topic from at least one topic as subtopics.

[0209] For example, if Topic A, Topic B, and Topic C are specified for the content to be analyzed, and Topic A has 5 quotations mapped to it, Topic B has 2 quotations mapped to it, and Topic C has 1 quotation mapped to it, the control unit (130) can determine Topic A, which has the most quotations, as the main topic. On the other hand, Topic B and Topic C, which have relatively fewer quotations, can be classified as subtopics.

[0210] As illustrated in FIG. 11d, the control unit (130) can use a large language model (141) to determine a main topic and a subtopic from a content to be analyzed (or a second content to be analyzed) from a user query (932b) received from a user terminal (10), and generate a topic analysis result (970) including a quotation mapped to each of the determined main topic and subtopic. At this time, the control unit (130) can use the large language model (141) to generate a topic analysis result in which main topic assignment information corresponding to the main topic (971) and subtopic assignment information corresponding to the subtopic (973) are distinguished and configured.

[0211] Specifically, the control unit (130) can generate a prompt (or a third prompt) requesting the determination of at least one of a main topic and a subtopic in the content to be analyzed (700b) based on the number of at least one quotation mapped to each of at least one topic, according to the user query (932b), using the prompt generation unit (131). Additionally, the control unit (130) can generate a prompt (or a third prompt) requesting the generation of topic analysis content corresponding to the user query (932b).

[0212] Furthermore, the control unit (130) can process the generated prompt (or third prompt) as input to the large language model (141). Furthermore, the control unit (130) can obtain topic analysis content from the large language model (141) that includes a topic analysis result (970) in which main topic assignment information corresponding to the main topic and sub-topic assignment information corresponding to the sub-topic are separated and included.

[0213] As an example, the large language model (141) can identify a specific main topic (e.g., “Strengthening supply chain cooperation and energy security at the summit”) among at least one topic based on the number of at least one quote mapped to each of at least one topic in the content to be analyzed (or, the second content to be analyzed). Additionally, the large language model (141) can extract at least one main quote (972) related to the main topic, wherein the main quote (972) may include at least one quote corresponding to the main topic in the content to be analyzed (700b).

[0214] Furthermore, the large language model (141) can determine a subtopic (973) by excluding a main topic from at least one topic, and generate a topic analysis result (970) including at least one of main topic assignment information including at least one main quote (972) mapped to the main topic (971) and subtopic assignment information including at least one sub quote (974) mapped to the subtopic (973).

[0215] The control unit (130) can obtain a topic analysis result (970) including at least one of main topic assignment information and sub-topic assignment information from a large language model (141). Furthermore, the control unit (130) can provide topic analysis content to a user terminal (10) including at least one of the topic analysis result (970) and response information generated based on the topic analysis result.

[0216] As described above, the quotation-based topic assignment method and system using a large language model according to the present invention can accurately assign topics to multiple sentences constituting the content to be analyzed based on quotations by utilizing a large language model. Through this, the accuracy and reliability of the topic analysis results can be improved by explicitly deriving the sentences or sentence segments that serve as the basis for topics in various forms of content (e.g., documents, articles, reviews, etc.).

[0217] Furthermore, the quotation-based topic assignment method and system using a large language model according to the present invention utilizes a large language model to extract quotations corresponding to at least one topic from the content to be analyzed and to assign topics on a quotation basis. Through this, the problem of topic contamination can be resolved when analyzing topics of the content to be analyzed, and semantic consistency within a single topic can be enhanced.

[0218] Furthermore, the quotation-based topic assignment method and system using a large language model according to the present invention can provide a topic analysis result including topic assignment information for content to be analyzed by combining topics and quotations. Through this, users can grasp the core topic of the content to be analyzed, the relationships between topics, and key quotations corresponding to each topic at a glance, and consequently, improve document comprehension and decision-making efficiency.

[0219] Meanwhile, the present invention described above can be implemented based on a quantum computer. The present invention implemented based on a quantum computer may include a qubit-based quantum processor and quantum memory, and may include software and hardware interfaces optimized for quantum computation.

[0220] Quantum processors in quantum computers utilize qubits to efficiently process complex operations through parallel computation, quantum entanglement, and quantum superposition, which cannot be performed by the binary bits of classical computers. Quantum processors process data using quantum gates and can provide exponential speed improvements for specific problems.

[0221] Meanwhile, the present invention described above can be implemented as a program that is executed by one or more processes on a computer and can be stored on a computer-readable medium (or recording medium).

[0222] Furthermore, the present invention described above can be implemented as computer-readable code or instructions on a medium on which a program is recorded. That is, the present invention can be provided in the form of a program.

[0223] Meanwhile, computer-readable media include all types of recording devices in which data that can be read by a computer system is stored. Examples of computer-readable media include HDD (Hard Disk Drive), SSD (Solid State Disk), SSD (Silicon Disk Drive), ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.

[0224] Furthermore, the computer-readable medium may be a server or cloud storage that includes a storage and is accessible to an electronic device via communication. In this case, the computer may download the program according to the present invention from the server or cloud storage via wired or wireless communication.

[0225] A computer program may reach the system (100) through various suitable transmission mechanisms. The transmission mechanism may be, for example, a computer-readable storage medium, a computer program product, a memory device, a recording medium such as a CD-ROM or DVD, or a product that tangibly embodies the computer program. The transmission mechanism may be a signal configured to reliably transmit the computer program through air or an electrical connection. The system (100) may propagate or transmit the computer program as a computer data signal.

[0226] Furthermore, references to 'computer-readable storage media,' 'computer program products,' 'computer programs embodied in a tangible form,' etc., or to 'controller,' 'computer,' 'processor,' etc., should be understood to include not only computers with various architectures such as single / multi-processor architectures and sequential (Von Neumann) / parallel architectures, but also specialized circuits such as Field-Programmable Gate Arrays (FPGAs), Application Specific Circuits (ASICs), signal processing units, and other devices. References to computer programs, instructions, code, etc., should be understood to include software for programmable processors or firmware, such as programmable content for hardware devices, whether it is instructions for a processor or configuration settings for a fixed-function device, gate array, or programmable logic device.

[0227] Furthermore, in the present invention, the computer described above is an electronic device equipped with a processor, namely a CPU (Central Processing Unit), and no special limitations are placed on its type.

[0228] Meanwhile, the above detailed description should not be interpreted restrictively in all respects but should be considered exemplary. The scope of the invention shall be determined by a reasonable interpretation of the appended claims, and all modifications within the equivalent scope of the invention are included within the scope of the invention.

Claims

1. Regarding methods performed by a computer, A step of specifying the content to be analyzed, consisting of at least one sentence; A step of specifying at least one topic related to at least one sentence constituting the content to be analyzed; A step of extracting at least one quotation corresponding to each of the at least one topic from the content to be analyzed using a large language model; and A quotation-based topic assignment method using a large language model, characterized by including the step of providing a topic analysis result for the content to be analyzed using the above topic and the above quotation.

2. In Paragraph 1, The step of specifying at least one topic above is, A step of generating a first prompt requesting the generation of at least one topic candidate composed of at least one of a Domain and a Category based on the above-mentioned content to be analyzed; A step of processing the first prompt as input to the large language model to obtain the at least one topic candidate from the large language model; and A quotation-based topic assignment method using a large language model, characterized by including the step of specifying the at least one topic related to the at least one sentence based on the at least one topic candidate.

3. In Paragraph 2, The step of extracting at least one quotation above is, A step of generating a second prompt requesting the extraction of the at least one quotation based on the above-mentioned content to be analyzed and the above-mentioned at least one topic candidate; and A quotation-based topic assignment method using a large language model, characterized by including the step of processing the second prompt as input to the large language model and extracting at least one quotation from the content to be analyzed from the large language model.

4. In Paragraph 3, The above second prompt is, A quotation-based topic assignment method using a large language model, characterized in that the large language model is configured to extract at least one word set consisting of at least one word that is semantically consecutive in at least one sentence as at least one quotation.

5. In Paragraph 3, A quotation-based topic assignment method using a large language model, characterized by further including the step of assigning a specific topic among the at least one topic to each of the at least one quotation using the large language model.

6. In Paragraph 5, The above second prompt is, A quotation-based topic assignment method using a large language model, characterized by being configured to map at least one specific topic among at least one topic to at least one specific quotation corresponding to the specific topic.

7. In Paragraph 6, The step of assigning the above specific topic is, The step of processing the above second prompt as input to the above large language model; and A quotation-based topic assignment method using a large language model, characterized in that the large language model includes the step of outputting at least one topic assignment information in which the specific topic and at least one specific quotation are mapped.

8. In Paragraph 7, The results of the above topic analysis are, A quotation-based topic assignment method using a large language model, characterized in that, based on the above-mentioned at least one topic assignment information, the above-mentioned at least one quotation included in the content to be analyzed is configured to be classified and displayed according to the above-mentioned at least one topic.

9. In Paragraph 7, The results of the above topic analysis are, A quotation-based topic assignment method using a large language model, characterized by including quotation reference information related to the location of at least one quotation included in at least one topic assignment information in the above-mentioned content to be analyzed.

10. In Paragraph 1, A step of generating topic analysis content from the analysis target content according to a user query received from a user terminal; and A quotation-based topic assignment method using a large language model, characterized by further including the step of providing the generated topic analysis content to the user terminal as a response to the user query.

11. In Paragraph 10, The above topic analysis content is, A quotation-based topic assignment method using a large language model, characterized by being configured to include at least one of the topic analysis result corresponding to the above user query and the response information to the above user query.

12. In Paragraph 11, In the step of generating the above topic analysis content, A step of generating a prompt requesting the generation of topic analysis content based on the above user query and the above analysis target content; and A quotation-based topic assignment method using a large language model, characterized by including the step of inputting the above prompt into the large language model and obtaining the above topic analysis content related to the user query from the large language model.

13. In Paragraph 12, In the step of generating the above prompt Collect user information corresponding to the user account logged into the above user terminal, and A quotation-based topic assignment method using a large language model, characterized by generating a prompt requesting the large language model to generate topic analysis content based on the above user information and the above user query.

14. In Paragraph 11, The above response information is, A quotation-based topic assignment method using a large language model, characterized by being configured to include summary information related to the user query extracted from the above topic analysis results.

15. In Paragraph 7, A quotation-based topic assignment method using a large language model, characterized by further including the step of determining a main topic and a subtopic among at least one topic related to the content to be analyzed.

16. In Paragraph 15, In the step of determining the main topic and subtopic above, A quotation-based topic assignment method using a large language model, characterized by determining the main topic among the at least one topics based on the number of at least one quotations mapped to each of the at least one topics.

17. In Paragraph 16, The results of the above topic analysis are, A quotation-based topic assignment method using a large language model, characterized in that main topic assignment information corresponding to the main topic and sub-topic assignment information corresponding to the sub-topic are configured separately.

18. In Paragraph 1, The above quote is, A quotation-based topic assignment method using a large language model, characterized by being at least one word set composed of at least one word that is semantically consecutive in at least one sentence.

19. Memory for storing instructions; and It includes at least one processor electrically connected to the memory, and When the above instructions are executed by the at least one processor, the at least one processor, Identifying the content to be analyzed, consisting of at least one sentence, and At least one topic related to the at least one sentence constituting the content subject to analysis is specified, and Using a large language model, at least one quotation corresponding to each of the at least one topic is extracted from the content to be analyzed, and A quotation-based topic assignment system using a large language model, characterized by providing a topic analysis result for the content to be analyzed using the above topic and the above quotation.

20. A program that is executed by one or more processes in an electronic device and stored on a computer-readable recording medium, The above program is, A step of specifying the content to be analyzed, consisting of at least one sentence; A step of specifying at least one topic related to at least one sentence constituting the content to be analyzed; A step of extracting at least one quotation corresponding to each of the at least one topic from the content to be analyzed using a large language model; and A program stored on a computer-readable recording medium characterized by including instructions that perform the step of providing a topic analysis result for the content to be analyzed using the above topic and the above quotation.

Citation Information

Patent Citations

  • Analysis apparatus and method for product trends and sale based on social big data

    KR1020160121132A

  • Electrochemical-active structure of nano-biohybrid actuator for enhancing movement and method for manufacturing thereof

    KR1020240058477A

  • System and method for determining and delivering breaking news utilizing social media

    US20160359790A1

  • Systems and methods for quote extraction

    US20170199932A1

  • Extracting quotes from customer reviews regarding collections of items

    US8700480B1