Information processing method, program, and information processing apparatus
The system uses large-scale language models to generate metadata and keywords for content, enabling the delivery of highly relevant advertisements by categorizing content and matching advertisements to these categories.
Patent Information
- Application Number
- JP2024123445
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2026-02-12
AI Technical Summary
Conventional metadata generation for content is inflexible and struggles to determine relevance between content and advertisements effectively.
An information processing system utilizing two large-scale language models to generate metadata and keywords for content, classify the content into categories based on similarity, and deliver advertisements matching those categories.
Enables the delivery of highly relevant advertisements by understanding the context of content and generating appropriate metadata and keywords, ensuring advertisements align closely with the content.
Smart Images

Figure 2026022082000001_ABST
Abstract
Description
[Technical Field]
[0001] The disclosed technology relates to an information processing method, a program, and an information processing device. [Background technology]
[0002] BACKGROUND ART Conventionally, a system is known that recognizes characters and sounds contained in video content such as news programs with high accuracy and automatically generates accurate metadata related to each video content (see, for example, Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Patent Publication No. 2021-12466 Summary of the Invention [Problem to be solved by the invention]
[0004] However, conventional technologies automatically generate metadata for content using machine learning, and the generation of metadata is dependent on a pre-trained model, making it difficult to generate flexible metadata. Furthermore, when using metadata to deliver advertisements, there is room for improvement in determining the relevance between content and advertisements.
[0005] Therefore, an object of the disclosed technology is to provide an information processing method, an information processing device, and an information processing system that enable the transmission of advertisements that are highly relevant to content. [Means for solving the problem]
[0006] In one embodiment of the disclosure, an information processing device performs an information processing method in which the information processing device inputs content and a request to generate metadata from the content into a first large-scale language model, obtains metadata for the content output from the first large-scale language model, inputs a request to generate keywords for each category of the content into a second large-scale language model, obtains the keywords for each category output from the second large-scale language model, classifies the content into at least one of the categories based on the similarity between the obtained keywords and the metadata, identifies advertisements that match the category into which the content is classified, and controls the transmission of the identified advertisements. [Effects of the Invention]
[0007] The disclosed technology makes it possible to send advertisements that are highly relevant to content. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a diagram illustrating an example of a configuration of an information processing system according to an embodiment. [Figure 2] FIG. 1 is a diagram illustrating an example of a configuration of an information processing device according to an embodiment. [Figure 3] FIG. 2 is a diagram illustrating an example of metadata of content according to an embodiment. [Figure 4] FIG. 10 is a diagram showing an example of keywords for each category according to an embodiment. [Figure 5] FIG. 10 illustrates an example of a metadata generation prompt according to one embodiment. [Figure 6] FIG. 10 illustrates an example of a keyword generation prompt according to one embodiment. [Figure 7] FIG. 2 is a diagram illustrating an advertising space and an advertising playlist according to an embodiment. [Figure 8] FIG. 1 illustrates an example of real-time processing according to an embodiment. [Figure 9]10 is a flowchart illustrating an example of a process related to advertisement transmission according to an embodiment. [Figure 10] 10 is a flowchart illustrating an example of a process related to content division according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0009] Preferred embodiments of the present disclosure will be described with reference to the accompanying drawings. In the drawings, components with the same reference numerals have the same or similar configurations.
[0010] An information processing system in the disclosed technology will be described below. In the disclosed technology, "broadcast" is a concept that includes television broadcast and radio broadcast. Furthermore, "program" includes television broadcast and radio broadcast programs that are broadcast in a predetermined broadcast time slot. Advertisements are widely publicized to the general public, and include advertisements broadcast via terrestrial waves, etc., and in some cases, advertisements distributed via the Internet. Television commercials and radio commercials are collectively referred to as commercials (CM), and commercials are a type of advertisement.
[0011] <Summary of the disclosed technology> The outline of the disclosed technology is that an information processing system of the disclosed technology uses a large-scale language model to detect objects, etc. contained in content (video / audio), and assigns metadata (which may be keywords, etc.) based on the detected objects, etc. to the content. The information processing system also uses the large-scale language model to generate keywords appropriate for each category of content. The information processing system classifies this content into categories having keywords that match the content's metadata, and controls the delivery of advertisements corresponding to the classified categories.
[0012] The information processing system described above, for example, makes it possible to obtain flexible and appropriate metadata for content by having a generation AI (Artificial Intelligence) generate metadata for content, and also makes it possible to obtain appropriate keywords that incorporate trends, new words, etc. by having the generation AI generate category keywords. This makes it possible to identify advertisements to be sent using categories that have keywords that match the appropriate metadata, and to send advertisements that are more relevant to the content.
[0013] Furthermore, when the content is video, conventional AI generally performs object detection on still images that break up the content (video / audio) into frames, making it difficult to properly extract the context of the video. However, by using the disclosed technology, it becomes possible to understand the context of the content (video / audio) and obtain metadata, thereby enabling the delivery of advertisements that are more suited to the content (video / audio).
[0014] [Embodiment] In the embodiment described below, a large-scale language model is appropriately utilized to enable the delivery of advertisements that are highly relevant to content. For example, an information processing system according to the embodiment inputs content and a prompt for generating metadata to a generation AI, and obtains the metadata generated by the generation AI. The information processing system obtains keywords for each category of the content generated by the generation AI, classifies the content into at least one category using the keywords and the content metadata, and controls the delivery of advertisements using the categories. Next, each configuration of the above-mentioned information processing system will be described.
[0015] <System> 1 is a diagram illustrating an example of a configuration of an information processing system 1 according to an embodiment of the present disclosure. In the example illustrated in Fig. 1, the information processing system 1 that implements the above-described processing may include an information processing device 10, a first server 20A, a first database 20B, a second server 30A, and a second database 20C.
[0016] The first server 20A and the first database 20B may be configured to be provided as a single service, and similarly, the second server 30A and the second database 30B may be configured to be provided as a single service. For example, the nth database stores a large amount of data used in a large-scale language model, and the nth server performs AI processing using the data stored in the nth database. Furthermore, the first server 20A, the second server 30A, the first database 20B, and the second database 30B may constitute a single service. For example, this service includes a service that can be provided on the cloud.
[0017] The devices in the information processing system 1 transmit and receive data to and from each other via a network N. Each server may be composed of multiple processing devices (which may include databases). Note that the "n" in the nth server and nth database is a number used to identify each server, and the number may be changed as needed.
[0018] The network N is composed of a wireless network or a wired network. Examples of the network include a mobile phone network, a PHS (Personal Handy-phone System) network, a wireless LAN (a local area network, including communication conforming to IEEE802.11 (so-called Wi-Fi (registered trademark))), 3G (3rd Generation), LTE (Long Term Evolution), 4G (4th Generation), 5G (5th Generation), WiMax (registered trademark), infrared communication, visible light communication, Bluetooth (registered trademark), a wired LAN, a telephone line, a power line communication network, a network conforming to IEEE1394, etc.
[0019] The information processing device 10 is configured by a server, a personal computer, or the like, and transmits requests to each server and receives responses from each server via the network N. For example, the information processing device 10 is a device within a broadcasting station, and transmits a metadata generation request together with content to the first server 20A and receives content metadata from the first server 20A in response to a user operation.
[0020] The information processing device 10 transmits a request for generating keywords indicating each category of content to the second server 30A, and receives keywords for each category from the second server 30A. The information processing device 10 may periodically transmit a keyword generation request to the second server 30A to update the keywords as appropriate. This allows the information processing device 10 to acquire keywords corresponding to trends and newly appearing words.
[0021] The information processing device 10 classifies the content into at least one category using the similarity between the metadata of the content and the keywords of each category. The information processing device 10 identifies advertisements for the content according to the classified category. For example, when there is an empty position in the advertisement frame for the content, the information processing device 10 identifies advertisements that are highly relevant to the content and allocates them to the empty position. At this time, the information processing device 10 may identify advertisements based on scores assigned to the content using predetermined data.
[0022] The information processing device 10 controls so that the advertisement specified for the content is sent when the content is broadcast. For example, the information processing device 10 may send identification information of the specified advertisement together with the position of the advertisement frame to an advertisement sending server, or if the information processing device 10 has an advertisement sending function, send the specified advertisement at the position of the advertisement frame. Note that "sending" includes broadcasting as a television advertisement and distributing as an internet advertisement.
[0023] The first server 20A configures a generation AI using, for example, a first large-scale language model stored in the first database 20B. The generation AI can be an existing generation AI. The information processing device 10 can send requests and receive responses via an API (Application Programming Interface) of the generation AI.
[0024] For example, the first server 20A generates metadata for the content based on the content and a prompt related to metadata generation input from the information processing device 10. The first server 20A outputs the generated metadata to the information processing device 10.
[0025] The second server 30A configures a generation AI using, for example, a second large-scale language model stored in the second database 30B. The information processing device 10 can send requests and receive responses via the API of the generation AI.
[0026] For example, the second server 30A generates keywords for each category based on a prompt for generating keywords for the content category from the information processing device 10. The second server 30A outputs the generated keywords to the information processing device 10.
[0027] The first server 20A and the second server 30A may be the same server, and the first database 20B and the second database 30B may be the same database.
[0028] <Configuration of information processing device> Among the devices in the information processing system 1, the components of the information processing device 10 that performs processing to identify advertisements highly relevant to content will be described below.
[0029] 2 is a diagram showing an example of the configuration of an information processing device 10 according to an embodiment of the present disclosure. The information processing device 10 includes one or more processors (CPUs) 110, one or more communication interfaces 120, a storage device 130, a user interface 150, and one or more communication buses 170 for interconnecting these components.
[0030] The user interface 150 is connected to a display and an input device (such as a keyboard and / or a mouse or some other pointing device), and has display and input functions.
[0031] The storage device 130 may be, for example, a high-speed random access memory such as a DRAM, an SRAM, or other random access solid-state storage device, or may be a non-volatile memory such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices, or may be a non-transitory computer-readable recording medium.
[0032] The storage device 130 stores data used by the information processing system 1. For example, the storage device 130 stores content, metadata of generated content, keywords for each category, information about advertising materials, information about advertising spaces, and the like.
[0033] Another example of memory 130 is one or more memory devices located remotely from processor 110. In one embodiment, memory 130 stores programs, modules, and data structures, or a subset thereof, that are executed by processor 110.
[0034] The processor 110 executes a program stored in the storage device 130 to configure a control unit 111 having the functions of the information processing device 10 described above. The control unit 111 configures a first request unit 112, a second request unit 113, a classification unit 114, an identification unit 115, a sending unit 116, and a calculation unit 117.
[0035] As described above, control unit 111 executes the processes related to generating metadata for content, generating keywords for each category of content, classifying content, identifying advertisements, and sending. In addition, control unit 111 has first request unit 112, second request unit 113, classification unit 114, identification unit 115, sending unit 116, and calculation unit 117 to execute the processes described above.
[0036] The first request unit 112 has a first input unit 112A and a first acquisition unit 112B. The first input unit 112A inputs a request to generate content and metadata from this content to the first server 20A that configures the first large-scale language model. This request includes, for example, a prompt to cause the generation AI to generate metadata related to objects detected by object detection or the like from the input content, metadata detected from audio or the like, and metadata detected from subtitles, captions, or the like.
[0037] In response to the prompt, the first acquisition unit 112B acquires metadata of content output from the first server 20A that constitutes the first large-scale language model. For example, the first acquisition unit 112B acquires metadata of content generated by a generation AI executed by the first server 20A. The number of metadata is not particularly limited, but a lower limit and / or an upper limit may be set in the prompt.
[0038] The second request unit 113 has a second input unit 113A and a second acquisition unit 113B. The second input unit 113A inputs a request to generate keywords for each content category to the second server 30A that configures the second large-scale language model. This request includes, for example, a prompt that causes the generation AI to generate appropriate keywords for each content category.
[0039] Examples of categories include travel, home, gardening (house and landscaping), beauty, health, banking, finance, etc. For each category, the prompt may include definitions and conditions for extracting keywords, keywords to exclude, output format, etc.
[0040] In response to the prompt, the second acquisition unit 113B acquires keywords for each category output from the second server 30A that constitutes the second large-scale language model. For example, the second acquisition unit 113B acquires at least one keyword for each category generated by the generation AI executed by the second server 30A. The number of keywords is not particularly limited, but a lower limit and / or an upper limit may be set in the prompt.
[0041] The classification unit 114 classifies the content into at least one of the categories based on the similarity between the keywords acquired by the second acquisition unit 113B and the metadata acquired by the first acquisition unit 112B. For example, the classification unit 114 may classify the content into the category with the largest average value of the similarity between each metadata and each keyword, or may have the generation AI perform the category classification. The classification unit 114 may use any method as long as it can classify the content using the metadata of the content generated by the generation AI and the keywords of each category.
[0042] The identification unit 115 identifies an advertisement that matches a category into which the content is classified. For example, when there is a vacant position in an advertisement slot associated with content that is a program, the identification unit 115 identifies an advertisement that corresponds to the category of the content as an advertisement to be allocated to the vacant position. As a specific example, the identification unit 115 may determine a match between the content category and an advertisement based on predetermined conditions such as genre and constraints, or may identify an advertisement to be allocated to the vacant position at the request of an advertiser or advertising agency.
[0043] The sending unit 116 controls the sending of the advertisement identified by the identification unit 115. For example, the sending unit 116 may output the position of the advertisement space and identification information of the identified advertisement to an external advertisement sending server. Furthermore, if the sending unit 116 has an advertisement sending function, it may control the sending of the identified advertisement according to the position of the advertisement space.
[0044] By using the above process and categories that have keywords that match metadata appropriate for the content, it becomes possible to identify the advertisement to be sent, making it possible to deliver advertisements that are more relevant to the content.
[0045] The second input unit 113A may input a keyword generation request to the second large-scale language model at a predetermined timing. For example, the predetermined timing may be every predetermined period of time, e.g., every few months. This allows the second input unit 113A to periodically automatically generate prompts and input them to the generation AI.
[0046] The second acquisition unit 113B may acquire updated keywords for each category from the second large-scale language model. For example, the second acquisition unit 113B may acquire keywords for each category in response to a keyword generation prompt that is periodically input to the generation AI. The second acquisition unit 113B updates the keywords for each category stored in the storage device 130 to the latest keywords.
[0047] By performing the above process, keywords for each content category are acquired each time a prompt is entered at a predetermined timing, allowing keywords that take into account the latest trends and newly emerging words to be acquired. As a result, it becomes possible to perform category classification using updated keywords for each category.
[0048] The first input unit 112A may input a request to generate metadata for content including videos, such as programs and news, to the generation AI provided by the first server 20A. In this case, the first acquisition unit 112B may acquire metadata that is in line with the context of the video, generated by a first large-scale language model to which content including the video has been input. For example, in the case of content including a video, the first acquisition unit 112B acquires metadata generated by a generation AI that uses the first large-scale language model. The prompt may request that metadata be generated in line with the context of the video. For example, the prompt may be such that metadata can be understood based on the connections between multiple images in the video.
[0049] Through the above processing, in the case of content that includes video, it is possible to determine the relevance of advertisements using metadata that is in line with the context of the video, making it possible to identify relevant advertisements based on the content of the content.
[0050] The calculation unit 117 calculates a score for the content based on factors that contribute to viewing, including at least one of stock prices, news, weather, time zone, and season. For example, the calculation unit 117 acquires information on other services that are linked with the control unit 111, such as stock prices, news, weather, etc., through API linkage with such services, or acquires the time zone and season based on the current date and time. The calculation unit 117 uses at least one of the acquired data to calculate a score using a predetermined calculation method.
[0051] The calculation unit 117 may change the data used to calculate the score depending on the category of the content. For example, when the category of the content is news, the calculation unit 117 may score the attention level of the news based on the number of news articles containing the metadata of the content. Furthermore, the calculation unit 117 may score the degree of compatibility between the metadata of the content and the acquired weather, time period, or season, or, if the metadata of the content includes economic terms, may calculate the score based on the ups and downs of the acquired stock prices.
[0052] The calculation unit 117 may calculate a score for content based on viewing history data including viewing ratings or non-specific viewing data, and / or viewing trend data including live streaming results or data related to content on SNS. For example, the calculation unit 117 calculates the score using a calculation method that gives a higher score the longer the viewing history, based on viewing ratings obtained from a company that surveys television viewing ratings, or viewing history data including non-specific viewing data that can be identified by an IP address or the like.
[0053] For example, calculation unit 117 acquires the distribution record of when content is / was live distributed, or acquires viewing trend data including comments (tweets, etc.) on content identified by the content title, metadata, etc. from a web service such as a social networking site. Calculation unit 117 calculates the score using a calculation method that increases the score when there is a large distribution record or when the viewing trend data indicates excitement (such as a large number of comments).
[0054] The calculation unit 117 may calculate a score related to the impression of the content, which includes at least the duration of stay of an object in the content, the telop size, or the type of source of appearance of a keyword. For example, the calculation unit 117 calculates a score that takes into account the impression by a predetermined calculation method related to the impression, using the duration of stay of an object detected by object detection or the like in the content, the telop size (or the character size relative to the screen size), and whether the keyword appeared in a telop or was detected and appeared by audio.
[0055] For example, the calculation unit 117 may calculate the score using a calculation method in which the score increases as the object stay time increases, the telop size increases, and the number of types of keyword appearance sources increases.
[0056] The above process makes it possible to calculate and assign a score to content. The calculated score may be used to identify advertisements, as described below. Note that the calculation unit 117 may use at least one of the scores described above, or may use a total score calculated by assigning weights to each score. Also, which score to use may be set by the user or based on the genre of the content.
[0057] The identification unit 115 may identify a matching advertisement using the category into which the content is classified and the score of the content. For example, the identification unit 115 identifies an arbitrary advertisement from advertisements assigned a category (genre) similar to the category into which the content is classified. For example, if the content is classified into the category of sports, the identification unit 115 identifies at least one advertisement from among advertisements whose genre is sports.
[0058] By performing the above process, it is possible to more appropriately identify advertisements that are related to the content by using the score calculated for the content.
[0059] The first input unit 112A may divide the content into predetermined units and input each of the divided predetermined units to the first large-scale language model. For example, the first input unit 112A may divide the content into predetermined units if the capacity or playback time of the content is equal to or greater than a threshold. The first input unit 112A may input each divided predetermined unit to the generation AI.
[0060] The first input unit 112A may input an instruction to divide the content into predetermined units into the metadata generation prompt, which enables the generation AI to divide the content into predetermined units and generate metadata for the content in the predetermined units.
[0061] (real-time processing) When the content is a live broadcast, the first request unit 112 may input the live broadcast video content to the generation AI while including in the metadata generation prompt an instruction to divide the video content into predetermined units. Alternatively, the first request unit 112 may divide the live broadcast video content into predetermined units as preprocessing and input each predetermined unit to the generation AI.
[0062] When the content is a live broadcast, the identification unit 115 identifies an advertisement to be allocated to a vacant position in the advertisement slot using metadata of a predetermined unit of content a predetermined time before the vacant position. This is because if an advertisement is identified based on the content immediately before the vacant position, there is a possibility that the metadata generation process by the generation AI, the category matching process, and the advertisement identification process may not be completed in time for the advertisement to be sent, so it is advisable to provide a buffer of a predetermined time before the advertisement is sent.
[0063] <Data example> Next, an example of data will be described. FIG. 3 is a diagram showing an example of metadata for content according to an embodiment of the disclosure. As shown in FIG. 3, the metadata generated by the generation AI is stored in association with the identification information of the content. For example, the first acquisition unit 112B stores the metadata acquired from the first server 20A in the storage device 130 in association with the identification information of the content input together with the metadata generation request. For example, metadata 1 "Cheers" and metadata 2 "Wine" are stored in association with the content "B1."
[0064] FIG. 4 is a diagram showing an example of keywords for each category according to an embodiment of the present disclosure. As shown in FIG. 4, the keywords generated by the generation AI are stored in association with the category identification information. For example, the second acquisition unit 113B stores the keywords acquired from the second server 30A in the storage device 130 in association with the category identification information input together with the keyword generation request. For example, keyword 1 "Hawaii," keyword 2 "Airplane," keyword 3 "Camping," etc. are stored in association with category "C1 (Travel)."
[0065] <Example of a prompt> Next, examples of prompts will be described. FIG. 5 is a diagram illustrating an example of a metadata generation prompt according to an embodiment of the present disclosure. As shown in FIG. 5, the prompt may include various instructions, such as an indication that the generation AI is a video content professional (SYSTEM_INSTRUCTION), an indication that the content should be analyzed to extract video information (TEMPLATE), analysis instructions (Analysis Instructions), an output format (Output Format), and instructions for extracting humans, animals, etc. (INSTRUCTIONS). When video content is divided, the start time or end time of each video content is associated. This enables the generation AI to generate metadata from the content in accordance with the above-described prompt.
[0066] FIG. 6 is a diagram illustrating an example of a keyword generation prompt according to an embodiment of the present disclosure. As shown in FIG. 6, the prompt may include various instructions such as a category name, a category definition, a condition for keyword generation, and an output keyword and score. The above prompt is generated for each category. This enables the generation AI to generate keywords for each category according to the above prompt.
[0067] <Advertising space and advertising playlist> FIG. 7 is a diagram illustrating an advertising space and an advertising playlist according to an embodiment of the present disclosure. The advertising space shown in FIG. 7 is a predetermined advertising broadcast time slot during which advertisements for products, services, etc. are broadcast. At least a broadcast date and time and a broadcast station are determined for each advertising space. In the example shown in FIG. 7, the advertising space is an advertising space for television commercials. However, this is not limited thereto, and advertisements broadcast using the advertising space may also be radio commercials.
[0068] Ad slot information, which is information about the ad slot, is associated with each ad slot. Ad slot information is general information about each ad slot and is provided by an ad slot provider, such as a broadcast station. The ad slot information includes at least the date and time when the advertisement will be broadcast in the ad slot (advertising broadcast time slot) and the broadcast station. It may also include the type of ad slot, the number of seconds of ad slot, the name of the program associated with the ad slot (specifically, the name of the program in which the ad slot is broadcast or immediately preceding the broadcast time), and the price (purchase price) for the ad slot. Here, the type of ad slot (genre) is information indicating the relationship between the ad slot and the program broadcast time, such as "PT (participating commercial: a time set within the program broadcast time)" and "SB (station break: the time between one program and the next)."
[0069] Furthermore, the advertisement space information may further include information other than the above, such as the broadcast region (area information), broadcast period, broadcast time slot (the time slot during a day when the advertisement is broadcast: Daypart), information about the genre of the program associated with the advertisement space and the performers of that program, information about events related to the program, and targets, which will be described later. However, these are merely examples and the information is not limited to these.
[0070] One ad slot can broadcast multiple ads. The order in which ads are broadcast within one ad slot is called the ad position. An ad position is also called "PIB (Position in Break)" and indicates the order in which an ad will be broadcast, for example, if four 15-second ads can be broadcast within one ad slot (number of positions = 4).
[0071] Furthermore, each ad space is assigned an ad space identification information (ad space ID) that uniquely identifies the ad space. This ad space ID makes it possible to identify one ad space from among many ad spaces.
[0072] Next, an advertisement playlist will be described with reference to Figure 7. An advertisement playlist is a list that is associated with an advertisement slot ID and includes an advertisement material ID that identifies the advertisement material to be broadcast in that advertisement slot, the position of the advertisement, and information about the number of seconds (length information) of the advertisement material. Figure 7 shows an example of an advertisement playlist with four positions, with the number of seconds information and advertisement material IDs listed in order of position. For example, guaranteed advertisements may be allocated by the commercial broadcasting system to positions 1 and 2, a programmatic advertisement that was won in the most recent auction to position 3, and an advertisement specified according to the above-mentioned content category to position 4.
[0073] <Example of real-time processing> Next, an example of real-time processing will be described. Fig. 8 is a diagram showing an example of real-time processing according to an embodiment of the present disclosure. As shown in Fig. 8, a live broadcast program 1 is divided into predetermined units, A, B, etc. The division of the program content may be performed on the generation AI side, or the divided content into predetermined units may be input to the generation AI as preprocessing.
[0074] In the example shown in Figure 8, the content used to determine which advertisement to allocate to position 1 in the ad slot is program 1A from a predetermined time ago, and the content used to determine which advertisement to allocate to position 2 is program 1B from a predetermined time ago. The predetermined time should be set taking into consideration the processing time required for the generation process by the generation AI and the advertisement identification process. This makes it possible to send highly relevant advertisements based on content that is neither too recent nor too far in the past, and that is appropriately timed to determine which advertisement to allocate.
[0075] <Operation description> Next, a description will be given of each operation of the information processing system 1. Fig. 9 is a flowchart showing an example of a process related to advertisement transmission according to an embodiment of the present disclosure. In the example shown in Fig. 9, each process is executed for one piece of content.
[0076] In step S102, the first input unit 112A of the information processing device 10 inputs a metadata generation request together with the content to the first large-scale language model. For example, the first input unit 112A inputs the content and a metadata generation prompt as shown in FIG. 5 to the generation AI implemented by the first server 20A.
[0077] In step S104, the first acquisition unit 112B of the information processing device 10 acquires the metadata generated by the generation AI. The first acquisition unit 112B stores the metadata in the storage device 130 in association with the identification information of the content.
[0078] In step S106, the second input unit 113A of the information processing device 10 inputs a keyword generation request to the second large-scale language model. For example, the second input unit 113A inputs a keyword generation prompt as shown in Fig. 6 to a generation AI implemented by the second server 30A. Note that the first server 20A and the second server 30A may be the same server, and the metadata generation prompt and the keyword generation prompt may be input to the same generation AI.
[0079] In step S108, the second acquisition unit 113B of the information processing device 10 acquires the keywords for each category generated by the generation AI. The second acquisition unit 113B stores the keywords in the storage device 130 in association with the identification information of the category.
[0080] In step S110, the classification unit 114 of the information processing device 10 classifies the content into at least one of the categories based on the similarity between the keywords acquired by the second acquisition unit 113B and the metadata acquired by the first acquisition unit 112B. The classification unit 114 may use any method as long as it can classify the content using the metadata of the content generated by the generation AI and the keywords for each category.
[0081] In step S112, the identifying unit 115 identifies an advertisement that matches the category that the content is classified into. For example, when there is a vacant position in an advertisement frame associated with content that is a program, the identifying unit 115 identifies an advertisement that corresponds to the category of this content as an advertisement to be allocated to this vacant position.
[0082] In step S114, the sending unit 116 controls the sending of the advertisement identified by the identification unit 115. For example, the sending unit 116 may output the position of the advertisement space and identification information of the identified advertisement to an external advertisement sending server. Furthermore, if the sending unit 116 has an advertisement sending function, it may control the sending of the identified advertisement according to the position of the advertisement space.
[0083] Through the above process, it is possible to identify advertisements to be sent using categories having keywords that match metadata appropriate for the content, and to deliver advertisements that are more relevant to the content. Note that the order of steps S102 to S104 and steps S106 to S108 does not matter. Furthermore, if the processes of steps S106 to S108 have been performed in advance, the processes shown in FIG. 9 may omit steps S106 to S108.
[0084] 10 is a flowchart illustrating an example of a process related to content division according to an embodiment of the present disclosure. In the example illustrated in FIG. 10, processes corresponding to S102 to S110 illustrated in FIG.
[0085] In step S202, the first input unit 112A divides the video content into predetermined units. The division process may be performed by a generation AI.
[0086] In step S204, first acquisition unit 112B acquires elements (people, animals, clothes, etc. detected by object detection) extracted from divided video content (hereinafter also referred to as "divided video").
[0087] In step S206, the first acquisition unit 112B assigns (associates) metadata corresponding to the acquired elements to the divided moving image.
[0088] In step S208, the classification unit 114 classifies the divided moving images into categories using the metadata assigned to the divided moving images and the keywords of each category, and assigns (associates) the classified categories to the divided moving images.
[0089] In step S210, the control unit 111 determines whether or not all divided moving images have been processed. If all divided moving images have been processed (step S210—YES), the process proceeds to step S212. If not all divided moving images have been processed (step S210—NO), the process returns to step S204, and the next divided moving image is processed.
[0090] In step S212, the classification unit 114 assigns (associates) a category to the entire video content based on the categories assigned to the divided videos. For example, the classification unit 114 may assign the most common category among the categories assigned to the divided videos to the video content.
[0091] The above process allows for appropriate processing of video content with long playback times, and enables the delivery of relevant advertisements in smaller increments. The process shown in Figure 10 can be applied to already generated video or audio content, and can also be applied to live broadcast content that is broadcast in real time.
[0092] The above-described embodiments are merely examples for explaining the disclosed technology, and are not intended to limit the disclosed technology to only the embodiments. The disclosed technology can be modified in various ways without departing from the spirit of the invention. For example, the processes of each device may be appropriately integrated, or the processes may be transferred to another device. [Explanation of symbols]
[0093] 1...information processing system, 10...information processing device, 20A...first server, 20B...first database, 30A...second server, 30B...second database, 110...processor, 130...memory, 111...control unit, 112...first request unit, 112A...first input unit, 112B...first acquisition unit, 113...second request unit, 113A...second input unit, 113B...second acquisition unit, 114...classification unit, 115...identification unit, 116...sending unit, 117...calculation unit, 130...storage device
Claims
1. The information processing device inputting content and a request to generate metadata from the content into a first large-scale language model; obtaining metadata for the content output from the first large-scale language model; inputting the request to generate keywords for each category of content into a second large-scale language model; obtaining keywords for each of the categories output from the second large-scale language model; classifying the content into at least one of the categories based on the similarity between the acquired keywords and the metadata; identifying advertisements that match the categories into which the content is classified; controlling delivery of the identified advertisements; An information processing method that performs the above.
2. Inputting the second large scale language model comprises: inputting the request into the second large-scale language model at a predetermined timing; Obtaining the keyword includes: The information processing method of claim 1 , further comprising obtaining updated keywords for each of the categories from the second large-scale language model.
3. Obtaining the metadata includes: The information processing method of claim 1 , further comprising obtaining metadata that is generated by the first large-scale language model to which content including a video is input, the metadata being in line with the context of the video.
4. The information processing method according to claim 1 , further comprising calculating a score for the content based on factors that contribute to viewing, including at least one of stock prices, news, weather, time of day, and season.
5. The information processing method of claim 1 further comprises calculating a score for the content based on viewing history data including viewing ratings or non-specific viewing data, and / or viewing trend data including live streaming performance or data related to the content on SNS.
6. The information processing method according to claim 1 , further comprising calculating a score relating to the impression of the content, the score including at least the duration of time spent on an object in the content, the telop size, or the type of source of a keyword.
7. The identifying step includes:
7. The information processing method according to claim 4, further comprising identifying a matching advertisement using a category into which the content is classified and a score of the content.
8. Inputting the first large scale language model comprises:
2. The information processing method according to claim 1, further comprising dividing the content into predetermined units and inputting each of the divided predetermined units.
9. In the information processing device, inputting content and a request to generate metadata from the content into a first large-scale language model; obtaining metadata for the content output from the first large-scale language model; inputting the request to generate keywords for each category of content into a second large-scale language model; obtaining keywords for each of the categories output from the second large-scale language model; classifying the content into at least one of the categories based on the similarity between the acquired keywords and the metadata; identifying advertisements that match the categories into which the content is classified; controlling delivery of the identified advertisements; A program that executes the following.
10. a first input unit that inputs content and a request to generate metadata from the content into a first large-scale language model; a first acquisition unit that acquires metadata of the content output from the first large-scale language model; a second input unit that inputs a request to generate keywords for each category of content into a second large-scale language model; a second acquisition unit that acquires keywords for each category output from the second large-scale language model; a classification unit that classifies the content into at least one of the categories based on the similarity between the acquired keywords and the metadata; an identification unit that identifies an advertisement that matches a category into which the content is classified; a sending unit that controls sending of the identified advertisement; An information processing device comprising:
Citation Information
Patent Citations
Metadata generation system, video content management system, and program
JP2021012466A