Movie metadata extraction system and method

The movie metadata extraction system addresses the limitations of structured data in movie recommendation systems by integrating structured and unstructured data through natural language processing, enabling more accurate and personalized recommendations.

WO2025135473A1PCT designated stage expired Publication Date: 2025-06-2602O CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/016914
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-18
Filing Date
2024-10-31
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing movie recommendation systems rely heavily on structured data and struggle to comprehensively analyze and understand unstructured data such as complex themes, emotional elements, and social contexts of movies.

Method used

A movie metadata extraction system and method that integrates structured and unstructured data through a data collection unit, preprocessing unit, and metadata extraction unit, utilizing natural language processing techniques to refine, convert, and quality-control unstructured data for keyword extraction and classification.

Benefits of technology

The system provides more accurate and personalized movie recommendations by capturing diverse user tastes and behaviors, improving content diversity, and enhancing user experience through a deeper understanding of movie metadata.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024016914_26062025_PF_FP_ABST
    Figure KR2024016914_26062025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed are a movie metadata extraction system and method. The movie metadata extraction system (method) according to the present invention comprises: collecting, by a data collection unit, structured data, which has a data form structured in a preset manner, and unstructured data, which has a data form other than the structured data; refining, by a preprocessing unit, the unstructured data through natural language processing technology, converting the refined unstructured data into an embedding vector, and performing quality control on the embedding vector on the basis of the structured data; and extracting, by a metadata extraction unit, a keyword by executing a specified prompt on the preprocessed unstructured data, and classifying at least one extracted keyword into respective specified categories through the natural language processing technology. According to the present invention, various requests and tastes of a user may be more accurately identified and satisfied by extracting metadata from structured data and unstructured data.
Need to check novelty before this filing date? Find Prior Art

Description

Movie metadata extraction system and method

[0001] The present invention relates to a technology for extracting movie metadata. More specifically, the present invention relates to a movie metadata extraction system and method for collecting movie-related data and extracting movie metadata.

[0002] Recently, technologies that utilize artificial intelligence (AI) to recommend movies have been widely utilized. These recommendations are primarily tailored to the user's preferences, viewing history, and other factors.

[0003] Recommending personalized movies like this requires building a movie database, which primarily utilizes structured data. Specifically, movie metadata is extracted using clearly defined structured data, such as genre, director, and actor, and then compiled into a database. Then, personalized movie recommendations are made through artificial intelligence learning based on user preferences, viewing history, and other factors.

[0004] However, while these movie recommendation systems work effectively with structured data, they have limitations in comprehensively analyzing and understanding unstructured data, such as complex themes, emotional elements, and social contexts of movies.

[0005] To overcome the limitations of these existing technologies, a new approach to analyzing unstructured natural language data is needed.

[0006] Unstructured data possesses the strength of containing diverse and in-depth information. Unstructured data, such as user reviews, social media posts, and expert ratings, provide a broad and multidimensional perspective on movies. Therefore, unstructured data can capture emotional, social, and cultural aspects of a movie that structured data cannot, making it valuable for movie recommendation systems.

[0007] However, one of the major challenges in utilizing unstructured data is the presence of low-quality data.

[0008] Therefore, the process of effectively selecting and filtering unstructured data is essential.

[0009] [Prior Art Literature]

[0010] [Patent Document]

[0011] (Patent Document 1) Document 1. Korean Intellectual Property Office Patent Registration No. 10-1639987, "Method and Device for Recommending Movies Based on Mixed Filtering"

[0012] Accordingly, the present invention has been made to solve the problems of the above-mentioned prior art, and an object of the present invention is to provide a movie metadata extraction system and method that can extract movie metadata by preset category using a plurality of models through purification, conversion, and quality control of collected data.

[0013] The movie metadata extraction system of the present invention for achieving the above purpose may preferably include a data collection unit that collects structured data having a data format structured in a preset manner and unstructured data having a data format other than the structured data; a preprocessing unit that refines the unstructured data using a natural language processing technique, converts the refined unstructured data into an embedding vector, and performs quality control on the embedding vector based on the structured data; and a metadata extraction unit that extracts keywords by executing a specified prompt on the preprocessed unstructured data, and classifies at least one or more extracted keywords into categories specified using a natural language processing technique.

[0014] As described above, the movie metadata extraction system and method according to the present invention has the following advantages.

[0015] 1. Traditional movie recommendation systems primarily rely on structured data (e.g., user ratings, purchase history), but these systems often fail to adequately reflect diverse user preferences and behaviors. Therefore, the present invention introduces a method that integrates diverse data sources, such as user interaction and behavior patterns, to address the limitations of structured data utilization, enabling more accurate prediction of user preferences and behaviors and providing personalized recommendations.

[0016] 2. Unstructured data (e.g., text, images, videos, etc. on social media) can help us gain a deeper understanding of user preferences and intentions. Therefore, the present invention utilizes unstructured data to gain a broader understanding of user preferences. Furthermore, by refining, transforming, and quality-controlling this unstructured data, we can understand and predict user needs in a more sophisticated and contextual manner.

[0017] 3. Content metadata, including existing movie metadata, often struggles with content diversity and recommendations for new content. The present invention enhances content diversity through analysis of unstructured data and improves content-related services by reflecting various aspects of new content. Furthermore, it can contribute to enhancing the user experience by reflecting diverse user preferences and the latest trends.

[0018] Figure 1 is a configuration diagram of a movie metadata extraction system according to one embodiment of the present invention.

[0019] Figure 2 is a configuration diagram of a data collection unit according to one embodiment of the present invention.

[0020] Figure 3 is a configuration diagram of a preprocessing unit according to one embodiment of the present invention.

[0021] Figure 4 is a configuration diagram of a metadata extraction unit according to one embodiment of the present invention.

[0022] Figure 5 is a flowchart of a movie metadata extraction method according to one embodiment of the present invention.

[0023] Figure 6 is a flowchart showing a data collection process according to one embodiment of the present invention.

[0024] Figure 7 is a flowchart showing a preprocessing process according to one embodiment of the present invention.

[0025] Figure 8 is a flowchart showing a metadata extraction process according to one embodiment of the present invention.

[0026] Hereinafter, the present invention will be described in detail with reference to preferred embodiments of the present invention and the accompanying drawings, on the premise that the same reference numerals in the drawings indicate the same components.

[0027] The present invention addresses the limitations of existing metadata extraction systems by complementing the limitations of structured data utilization and leveraging the advantages of unstructured data. In other words, the present invention introduces integrated utilization and analysis technology for structured and unstructured data to overcome the limitations of existing metadata extraction systems and improve the user experience. This can play a crucial role in more accurately identifying and satisfying users' diverse needs and preferences.

[0028] To this end, the present invention aims to extract integrated, AI-based movie metadata using KeyBERT, LLM (ChatGPT API), and KMDB data. This includes the process and methodology for selecting high-quality data from unstructured data and extracting and classifying movie-related keywords.

[0029] In this way, the present invention generates and classifies movie metadata based on natural language text extracted from unstructured data, particularly user reviews, social media, and expert evaluations. This provides a deeper and more multifaceted understanding of movies, enabling more accurate and personalized recommendations for users. This is achieved by integrating various sources, such as KeyBERT, LLM (ChatGPT API), and KMDB data, to capture the complex characteristics of movies and effectively apply them to a recommendation system. Furthermore, the process of effectively selecting and filtering unstructured data is essential, and natural language processing technology plays a crucial role in this process. The present invention utilizes these technologies to improve the quality of unstructured data, thereby enhancing the overall quality of recommendation services. This allows for more accurate and personalized recommendations to users.

[0030] Hereinafter, an example of implementing the movie metadata extraction system and method of the present invention will be described through a specific embodiment.

[0031] Figure 1 is a configuration diagram of a movie metadata extraction system according to one embodiment of the present invention.

[0032] Referring to FIG. 1, the movie metadata extraction system of the present invention includes a data collection unit (1) that collects structured data having a data format structured in a preset manner and unstructured data having a data format other than the structured data, a preprocessing unit (2) that refines the unstructured data through a natural language processing technique, converts the refined unstructured data into an embedding vector, and performs quality control on the embedding vector based on the structured data, and a metadata extraction unit (3) that extracts keywords by executing a specified prompt on the preprocessed unstructured data and classifies at least one or more extracted keywords into categories specified through a natural language processing technique.

[0033] Here, it's preferable to use highly reliable, publicly available data as structured data. Meanwhile, unstructured data refers to data other than structured data, such as data obtained from social media.

[0034] Meanwhile, preprocessing includes purification process, conversion process and quality control process.

[0035] The purification process involves removing profanity and slang from unstructured data.

[0036] Additionally, the transformation process involves converting unstructured data into an analyzable format, such as natural language processing for text analysis.

[0037] Additionally, the quality control process includes selecting unstructured data whose similarity to structured data is greater than a set value to remove unnecessary low-quality data.

[0038] Meanwhile, metadata extraction involves extracting metadata based on data type and characteristics, and categorizing the extracted metadata. Here, the semantic similarity of each keyword is used to categorize the data into four categories: atmosphere, plot, visual aesthetics, and music. Furthermore, keywords from Keybert or KMDB can undergo additional classification processes.

[0039] Here, the process of designing a database schema for classified metadata and databaseization may be further included. This process may include designing a database schema capable of efficiently storing and retrieving metadata, and selecting a model that considers the data format among various modeling methods, such as relational databases, NoSQL databases, and graph databases.

[0040] A database for storing metadata can be constructed based on a DB schema designed in this manner. In other words, a metadata database can be constructed.

[0041] Finally, it may include establishing an automated scheduling policy to periodically perform the process of integrating data collected from various sources, extracting metadata, and inserting it into a database.

[0042] Figure 2 is a configuration diagram of a data collection unit according to one embodiment of the present invention.

[0043] Referring to FIG. 2, the data collection unit (1) of the present invention includes a scheduler (11) that establishes a crawling policy in response to a new movie release cycle, a structured data collection unit (12) that collects structured data in the form of structured data from a public movie database or spreadsheet, and an unstructured data collection unit (13) that collects unstructured data in the form of unstructured data from social media posts, videos, music, and reviews.

[0044] Here, structured data includes information such as the movie's plot, rating, country of production, actors, and director, and can be collected from information provided by movie databases such as KMDB (Korean Movie Database).

[0045] Additionally, unstructured data includes information such as movie comments, images, videos, and text, and information can be collected through social media, etc.

[0046] Collected structured and unstructured data can be integrated into an integrated database.

[0047] Figure 3 is a configuration diagram of a preprocessing unit according to one embodiment of the present invention.

[0048] Referring to FIG. 3, the preprocessing unit (2) of the present invention includes a purification unit (21) that removes vulgar language, slang, and grammatical errors from unstructured data through data cleansing, a transformation unit (22) that divides the purified unstructured data into tokens, which are basic units of natural language processing, converts the divided tokens into unique IDs, pads the unique IDs so that they have the length of a set input sequence, and embeds the padded data, and a quality control unit (23) that measures the cosine similarity between a structured embedding vector obtained by inputting structured data into the transformation unit (22) and an unstructured embedding vector output from the transformation unit (22), and selects unstructured data having a similarity greater than a set value.

[0049] Figure 4 is a configuration diagram of a metadata extraction unit according to one embodiment of the present invention.

[0050] Referring to FIG. 4, the metadata extraction unit (3) of the present invention includes a keyword extraction unit (31) that extracts keywords by executing a plurality of artificial intelligence models and a preset category-specific prompt from a movie database on preprocessed and selected unstructured data, and a metadata classification unit (32) that classifies at least one extracted keyword by category through natural language processing technology.

[0051] At this time, the keyword extraction method can be carried out in four ways: extraction using Keybert and structured and unstructured data about the movie; extraction using LLM and structured and unstructured data about the movie; extraction using LLM's own knowledge by providing only the title and minimal information (title, production year) about the movie; and utilization of keywords manually written in KMDB. In addition, all keywords extracted in this way are integrated and classified into four categories: atmosphere, plot, visual beauty, and music.

[0052] Here, the metadata classification unit (32) may further include an additional classification unit (33) including an embedding unit (331) that converts each category text to be classified into a category embedding vector by describing the main keywords extracted for keywords that are not classified by category in the metadata classification process into sentences, and converts them into sentence embedding vectors, a similarity measurement unit (332) that measures the similarity between the sentence embedding vector and the category embedding vector, and a metadata selection unit (333) that selects keywords having a similarity greater than a set value as metadata.

[0053] This is to classify keywords into each category using natural language processing technology for keywords that have not yet been classified by category, such as Keybert and KMDB.

[0054] Then, the movie metadata extraction method of the present invention using the system configured as described above will be described.

[0055] Figure 5 is a flowchart of a movie metadata extraction method according to one embodiment of the present invention.

[0056] Here, the movie metadata extraction method according to one embodiment of the present invention is operated by linking each component of the above-described system, and each component of the system can be realized by operation between a processor and a memory.

[0057] Referring to FIG. 5, the movie metadata extraction method of the present invention includes a step of collecting structured data having a data format structured in a preset manner and unstructured data having a data format other than the structured data (S1), a step of refining the unstructured data using a natural language processing technology, converting the refined unstructured data into an embedding vector, and performing quality control on the embedding vector based on the structured data (S2), and a step of executing a designated prompt on the preprocessed unstructured data to extract keywords and classifying at least one or more extracted keywords into a designated category using a natural language processing technology (S3).

[0058] Additionally, the process of designing a database schema for classified metadata and databaseization may be further expanded. This process may include designing a database schema capable of efficiently storing and retrieving metadata, and selecting a model that considers the data format among various modeling methods, such as relational databases, NoSQL databases, and graph databases.

[0059] A database for storing metadata can be constructed based on a DB schema designed in this manner. In other words, a metadata database can be constructed.

[0060] Additionally, it may include establishing an automated scheduling policy to periodically perform the process of integrating data collected from various sources, extracting metadata, and inserting it into a database.

[0061] Figure 6 is a flowchart showing a data collection process according to one embodiment of the present invention.

[0062] Referring to FIG. 6, the data collection process of the present invention includes a step of establishing a crawling policy in response to a new movie release cycle (S11), a step of collecting structured data in the form of structured data from a public movie database or spreadsheet, and a step of collecting unstructured data in the form of unstructured data from social media posts, videos, music, and reviews (S12).

[0063] Here, structured data includes information such as the movie's plot, rating, country of production, actors, and director, and can be collected from information provided by movie databases such as KMDB (Korean Movie Database).

[0064] Additionally, unstructured data includes information such as movie comments, images, videos, and text, and information can be collected through social media, etc.

[0065] Collected structured and unstructured data can be integrated into an integrated database.

[0066] Figure 7 is a flowchart showing a preprocessing process according to one embodiment of the present invention.

[0067] Referring to FIG. 7, the preprocessing process of the present invention includes a step of removing profanity, slang, and grammatical errors through data cleansing for unstructured data (S21), a step of dividing the refined unstructured data into tokens, which are basic units of natural language processing, converting the divided tokens into unique IDs, padding them so that the unique IDs have the length of a set input sequence, and embedding the padded data (S22), and a step of measuring the cosine similarity between a structured embedding vector obtained by inputting structured data into a conversion unit (22) and an unstructured embedding vector output from the conversion unit (22), and selecting unstructured data having a similarity greater than a set value (S23).

[0068] Figure 8 is a flowchart showing a metadata extraction process according to one embodiment of the present invention.

[0069] Referring to FIG. 8, the metadata extraction process of the present invention includes a step of extracting keywords by executing a predetermined category-specific prompt from a plurality of artificial intelligence models and a movie database on preprocessed and selected unstructured data (S31), and a step of classifying at least one extracted keyword by category using natural language processing technology (S32).

[0070] At this time, the keyword extraction method can be carried out in four ways: extraction using Keybert and structured and unstructured data about the movie; extraction using LLM and structured and unstructured data about the movie; extraction using LLM's own knowledge by providing only the title and minimal information (title, production year) about the movie; and utilization of keywords manually written in KMDB. In addition, all keywords extracted in this way are integrated and classified into four categories: atmosphere, plot, visual beauty, and music.

[0071] Here, the classification step may further include a step of converting each category text to be classified into a category embedding vector by describing the main keywords extracted for the keywords in a sentence when there are keywords that are not classified by category in the metadata classification process (S33), a step of converting the sentence embedding vector into a sentence embedding vector, and a step of measuring the similarity between the sentence embedding vector and the category embedding vector (S35), and a step of selecting keywords having a similarity greater than a set value as metadata (S36).

[0072] This is to classify keywords into each category using natural language processing technology for keywords that have not yet been classified by category, such as Keybert and KMDB.

[0073] [Example]

[0074] Data collection

[0075] [Structured data]

[0076] Structured data refers to structured data formats that can be easily found in databases or spreadsheets. Structured data related to movies can include the following elements:

[0077] Title, release year, genre, director, lead actors, running time, country of production, language, budget, rating, rating...

[0078] This structured data provides fundamental, reliable, and essential information about movies and can be utilized for various purposes, including data analysis, movie recommendation systems, and market research. The present invention utilizes the APIs of Kofic and KMDB, Korean movie databases, and TMDB, an international movie database, to collect structured data.

[0079] Structured data includes basic information about each film, such as the year of production, director, and genre.

[0080] [Unstructured data]

[0081] In a similar vein to structured data, unstructured data refers to data that exists in an unstructured form. In the context of movies, unstructured data may include the following elements:

[0082] Reviews and ratings, movie trailers and clips, photos and posters, social media posts, interviews and scripts, audio tracks and music...

[0083] This unstructured data exists in various forms and formats, and analyzing and understanding it requires advanced data processing technologies and analysis methods. In this example, the unstructured data targeted are various movie synopses and reviews, and texts written in Korean, such as those found on Naver and Watchapedia.

[0084] Unstructured data provides various aspects of information such as consumer ratings and reviews for each movie.

[0085]

[0086] Preprocessing

[0087] [Data Cleansing]

[0088] Unstructured data may contain elements that degrade the quality of the text, such as repeated postings of the same review, text data that lacks useful information, or the use of profanity or grammatical errors. Therefore, natural language processing techniques, such as data cleansing, are used to remove these elements.

[0089] [Original example sample text]

[0090] "Hmm... I thought it was a movie for elementary school kids by looking at the poster... Even the overacting is trash, so it's not light."

[0091] This sample is excluded because it contains profanity and slang and does not contain useful information. In this case, a dictionary of Korean profanity and slang is built in advance, and any words containing these words are automatically excluded from the collected data.

[0092] It performs the task of refining and excluding low-quality data, which is one of the causes of deteriorating model performance.

[0093] [Data Conversion]

[0094] The input text for the Bert model is divided into tokens, which are defined as the minimum unit of data required for the model's input. The data conversion process is as follows.

[0095] 1. Tokenization: Input text can be sentences or documents, and tokenization is performed based on words. In natural language processing, tokenization is the process of breaking text into tokens, the smallest units that make up text. It refers to the process of breaking text into pieces that a model can understand and process.

[0096] 2. Adding special tokens: A CLS (classification) token is added to the beginning of each sentence for classification purposes, and a SEP (separation) token is added to the end. These special tokens are used to capture representative features, such as the meaning and context of the entire sentence. Examples are as follows.

[0097] [Examples of Special Token Usage]

[0098] Sentences to embed using Bert: The cat is sleeping. The dog is barking.

[0099] Sentences with special tokens: [CLS] The cat is sleeping. [SEP] The dog is barking. [SEP]

[0100] 3. Token ID Assignment: Each token is converted to a corresponding unique ID in the dictionary.

[0101] 4. Padding: Bert requires input of a fixed length, so all input sequences are padded to a maximum length. Short sequences are lengthened by adding PAD tokens. Padding is a technique used to make text data a fixed length, similar to special tokens, but with a different purpose.

[0102] [Example of using padding]

[0103] The sentence we want to embed using Bert: The cat is sleeping

[0104] Sentence using padding tokens: The cat is sleeping. [PAD][PAD][PAD][PAD]

[0105] 5. Utilizing Attention Masks: Bert uses attention masks to ignore padded data. The actual data is masked with "1" and the padded data with "0."

[0106] 6. Segmented Embedding: When two sentences are input, segmented embedding vectors are added to distinguish each sentence. While the application is similar to the SEP token, their purpose is different. The SEP token is used to indicate the end of a sentence, while the segmented embedding is used to distinguish two or more sentences or text fragments. This helps the model recognize that the input text fragments come from different sentences.

[0107] In other words, data transformation can be defined as the task of embedding data consisting of characters in a way that the model can understand.

[0108] [Data Quality Management]

[0109] Unstructured data, in the form of unstructured text, provides a wealth of information about a given product. However, due to its lack of standards and restrictions, it can also indiscriminately contain profanity, vulgarity, and unrelated words. Therefore, to improve the efficiency and performance of data analysis-based services and models, it is necessary to either eliminate or refine low-quality data. In this invention, unstructured data is compared with highly reliable product information, similar to structured data, and its quality is expressed numerically based on similarity. Furthermore, by selecting unstructured data above a certain threshold, overall data quality is secured.

[0110] [Data Quality Management Process: Similarity Verification]

[0111] The present invention utilizes movie synopses provided in structured data as highly reliable information on movie meta, and compares them with unstructured data such as reviews and blog posts to evaluate the quality of the unstructured data. The evaluation process is as follows.

[0112] 1. Structured and unstructured data are embedded using the "jhgan / ko-sroberta-multitask" natural language processing model, a Bert-based model.

[0113] 2. Measure the cosine similarity between the embedding vectors of structured and unstructured data.

[0114] 3. Based on the similarity measured between 0 and 1, as an example, select non-standard data with a similarity score of 0.7 or higher.

[0115]

[0116] Data Integration

[0117] An integrated database is built by utilizing the collected structured and unstructured data in one database.

[0118] An integrated database was built with structured and unstructured data to be used for meta extraction.

[0119]

[0120] Metadata extraction

[0121] Meta-extraction algorithms utilize various models, generative AI models, and handwritten keywords provided by KMDB. Keywords are extracted and implemented by category to reflect various aspects of movie content. For example, keyword categories could include atmosphere, plot, visual aesthetics, and music. The introduction of a keyword classification system is essential for systematically and efficiently managing large-scale keyword data. During the process of establishing this system, the types and number of categories are determined through numerous experiments and extraction trials. This process is part of an effort to achieve optimal results. Each category is designed to best reflect the characteristics of the data and enable analysts to understand and utilize the data more deeply.

[0122] [Algorithm and System Design]

[0123] In this example, two models (Keybert, LLM) and one database (KMDB) are used as meta-extraction algorithms, and four extraction methods are performed. The categories, which are the criteria for classifying each keyword, are defined as atmosphere, plot, visual beauty, and music. In summary, there are four ways to extract keywords: one method that extracts using Keybert and structured and unstructured data about movies; one method that extracts using LLM and structured and unstructured data about movies; one method that extracts by providing only the title and minimal information (title, production year) about the movie using LLM's own knowledge; and one method that utilizes manually written keywords in KMDB. In addition, all keywords extracted in this way are integrated and classified into four categories: atmosphere, plot, visual beauty, and music.

[0124] 1-1. Keyword Extraction Using Plots in Keybert: Keybert is a Bert-based keyword extraction model specialized for keyword extraction. When text is input, it measures the importance of keywords in the document and extracts them based on this information. Importance is measured by considering the frequency and location of keywords within each document.

[0125] [Example of Keybert output results]

[0126] Movie: Top Gun Maverick

[0127] Keywords printed: action package, nostalgia, mentoring, adventure and challenge, realistic flight scenes, visual spectacle, iconic soundtrack, intense action music

[0128] 1-2. Keyword Extraction Using Plots in LLM: Using Chatgpt version 4, currently recognized as the most powerful generative AI for natural language processing, we apply prompt engineering to the model, inputting the plot as input and extracting keywords related to the plot. Each keyword is output in a form classified into the four categories defined above.

[0129] [Example of output results using LLM plot]

[0130] Movie: Top Gun Maverick

[0131] Mood: Action-packed, nostalgic

[0132] Plot: Mentoring, Adventure, and Challenge

[0133] Visual Beauty: Realistic flight scenes, visual spectacle

[0134] Music: Iconic soundtrack, intense action music

[0135] [Prompt engineering]

[0136] It is a technique used in natural language processing, and is a technology that induces the generation of a desired output by designing an input (prompt) suitable for a specific AI model.

[0137] Prompt engineering is a technique for structuring input to achieve the desired output from large-scale language models, such as LLM. It involves inputting example samples together. Shot_1_input and Shot_1_output are examples, and Main_input corresponds to the question the LLM is supposed to answer.

[0138] 1-3. Keyword extraction using LLM's own knowledge without providing plot information: LLM learns from a large amount of data on the web and has a built-in understanding of known movies. Therefore, it can extract keywords for specific movies without providing specific information about them, and this information is used to extract keywords. This model also provides output classified into four categories.

[0139] [Example of output results without LLM plot]

[0140] Movie: Top Gun Maverick

[0141] Atmosphere: heroism, camaraderie, and friendship

[0142] Plot: Past and Present, Leadership and Responsibility

[0143] Visual Beauty: High-definition filming, dynamic camera work

[0144] Music: Harmony of music and video, various genres

[0145] 1-4. Utilizing KMDB Manual Keywords

[0146] Currently, KMDB has manually created keywords for each movie, and these keywords are used as a source for the metadata database.

[0147] [KMDB Manual Keyword Examples]

[0148] Movie: Top Gun Maverick

[0149] Keyword: instructor, fighter, pilot, training school, re-release

[0150] Keywords extracted from each model are integrated and used as meta information about the content.

[0151] Meanwhile, for keywords that have not yet been classified by category, such as Keybert and KMDB, natural language processing technology is used to classify the keywords into each category.

[0152] Machine Learning and Natural Language Processing Technologies

[0153] For the keywords extracted above, those from LLM are pre-categorized and extracted, but those from Keybert and KMDB are not categorized and require additional classification. Therefore, we utilize the "Ko-Sroberta-Multitask" category classification model, a Bert base, and classify each keyword into four categories: atmosphere, plot, visual beauty, and music, based on their semantic similarity. To achieve semantically effective classification, we utilize LLM to describe each keyword in a single sentence, and then use the corresponding texts together. The process is as follows.

[0154] 1. Enter the keywords to be classified as input to LLM and describe them in one sentence.

[0155] Example) Keywords to classify: Love --> "Love" is one of the important aspects of human emotions, and is often treated as a theme in movies, and is an important meta-topic that explores the complex emotions of humans through various relationships and emotional expressions.

[0156] 2. Embed the keywords to be classified and the sentences defined in LLM.

[0157] Example) Sentence to embed: "Love, love is one of the important aspects of human emotion, and is often treated as a theme in movies, exploring the complexities of human emotions through various relationships and emotional expressions."

[0158] 3. Embed text for each category (atmosphere, plot, visual beauty, music).

[0159] 4. Measure the cosine similarity between the embedding vectors from step 2 and the embedding vectors from step 3.

[0160] 5. Classify each keyword into the category with the highest similarity.

[0161]

[0162] DB schema design for classified metadata and DBization

[0163] [Metadatabase implementation]

[0164] The metadata classified by each category is built together with the corresponding movie and structured data to facilitate service utilization. The designed metadata database consists of various information about movie content. In addition to structured information such as the title, director, production year, country, and genre of the movie, it also contains various metadata such as keywords extracted from unstructured data, keyword similarity values, keyword extraction method, and data types used for keyword extraction. This metadata can be managed not only for a content similarity-based recommendation system that utilizes various information, but also for establishing and classifying a clustering or classification system by utilizing each characteristic. In addition, the extracted keywords can be provided together with each content on the UI / UX screen, so that they can be used to provide basic information about the content.

[0165] The meta database is a database that utilizes metadata built using data in the integrated database.

[0166]

[0167] Establishing a data collection and extraction scheduling policy

[0168] Establish scheduling policies to ensure that data is collected weekly, daily, and monthly, metadata is generated, and stored in the metadata database. For structured and unstructured movie data, crawling policies are established based on the release cycle of new movies.

Claims

1. A data collection unit that collects structured data having a structured data format in a preset manner and unstructured data having a data format other than the structured data; A preprocessing unit that refines the above unstructured data using natural language processing technology, converts the refined unstructured data into an embedding vector, and performs quality control on the embedding vector based on the structured data; and A movie metadata extraction system comprising a metadata extraction unit that extracts keywords by executing a specified prompt on preprocessed unstructured data and classifies at least one extracted keyword into a specified category using natural language processing technology.

2. In paragraph 1, The above data collection unit, A scheduler that establishes crawling policies in response to new movie release cycles; A structured data collection unit that collects structured data in the form of structured data from a public movie database or spreadsheet; and A movie metadata extraction system comprising an unstructured data collection unit for collecting unstructured data in the form of unstructured data from social media posts, videos, music, and reviews.

3. In paragraph 2, The above preprocessing unit, A purification unit that removes profanity, slang, and grammatical errors through data cleansing of the above unstructured data; A conversion unit that divides refined unstructured data into tokens, which are basic units of natural language processing, converts the divided tokens into unique IDs, pads the unique IDs so that they have the length of the set input sequence, and embeds the padded data; and A movie metadata extraction system including a quality management unit that measures the cosine similarity between a structured embedding vector obtained by inputting the structured data into the transformation unit and an unstructured embedding vector output from the transformation unit, and selects unstructured data having a similarity greater than a set value.

4. In paragraph 3, The above metadata extraction unit, A keyword extraction unit that extracts keywords by executing a number of artificial intelligence models and a preset category-specific prompt from a movie database on preprocessed and selected unstructured data; and A movie metadata extraction system including a metadata classification unit that classifies at least one extracted keyword into categories using natural language processing technology.

5. In paragraph 4, The above metadata classification section is, In the metadata classification process, the main keywords extracted for keywords that are not classified by the above categories are described as sentences and converted into sentence embedding vectors, and the embedding section converts each category text to be classified into a category embedding vector. A similarity measurement unit that measures the similarity between sentence embedding vectors and category embedding vectors, A movie metadata extraction system further comprising an additional classification unit including a metadata selection unit for selecting keywords having a similarity greater than a set value as metadata.

Citation Information

Patent Citations

  • Metadata extraction device and method therefor

    JP2010102668A

  • System and method for extraction performance improvement of unstructured text

    KR101644429B1

  • Social data analysis system for contents recommedation

    KR1020150096024A

  • System for providing multi-parameter analysis based commercial service using influencer matching to company

    KR102155342B1

  • System for recommending related data based on similarity and method thereof

    KR102411081B1