Information processing device, information processing method, and information processing program

The information processing device generates and adds contextually relevant captions to moving images by analyzing explanatory information, addressing the limitations of conventional voice recognition-based methods.

JP2026073824APending Publication Date: 2026-05-01LY CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
LY CORP
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Conventional techniques struggle to generate and add appropriate telops to the content of moving images, often relying solely on voice recognition without considering the context or content of the video.

Method used

An information processing device that acquires explanatory information about a moving image, generates captions based on this information using a trained model, and adds them to the video, considering factors like location, audio, facial expressions, and user preferences.

Benefits of technology

Enables the generation and addition of contextually appropriate captions to videos, enhancing user convenience by accurately representing the video's content, emotions, and situations, thus improving the video editing process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026073824000001_ABST
    Figure 2026073824000001_ABST
Patent Text Reader

Abstract

Generate and add appropriate captions to the content of the video. [Solution] The information processing device according to the present invention is characterized by comprising: an acquisition unit that acquires explanatory information describing the content of a video; a generation unit that generates captions to be added to the video (for example, captions indicating the subject of the video, captions indicating the state of mind of the subject of the video, captions indicating the situation of the subject of the video, etc.) based on the explanatory information acquired by the acquisition unit; and an application unit that adds the captions generated by the generation unit to the video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0004] , , ,

[0006] , ,

[0005] , , ,

[0007] , , ,

[0001] The present invention relates to an information processing apparatus, an information processing method, and an information processing program.

Background Art

[0002] Conventionally, techniques related to image processing performed by recognizing objects included in an image have been provided. As an example of such a technique, a technique of adding a telop using a voice detection result in a moving image to the moving image is known.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, with the above-described technique, it is not always possible to generate and add an appropriate telop to the content of the moving image.

[0005] For example, with the above-described technique, it only adds text data based on voice recognition from the moving image, and it is not always possible to generate and add an appropriate telop to the content of the moving image.

[0006] The present application has been made in view of the above, and an object thereof is to generate and add an appropriate telop to the content of a moving image.

Means for Solving the Problems

[0007] The information processing device according to the present application is characterized by comprising: an acquisition unit that acquires explanatory information describing the content of a moving image; a generation unit that generates a caption to be added to the moving image based on the explanatory information acquired by the acquisition unit; and an application unit that adds the caption generated by the generation unit to the moving image. [Effects of the Invention]

[0008] According to one embodiment, the effect is achieved that appropriate captions can be generated and added to the content of a video. [Brief explanation of the drawing]

[0009] [Figure 1] Figure 1 shows an example of information processing according to the embodiment. [Figure 2] Figure 2 shows an example of the configuration of the information processing device 10 according to the embodiment. [Figure 3] Figure 3 shows an example of a video information database 31. [Figure 4] Figure 4 is a flowchart showing an example of the information processing procedure according to the embodiment. [Figure 5] Figure 5 is a hardware configuration diagram showing an example of a computer that implements the functions of the information processing device 10. [Modes for carrying out the invention]

[0010] The following describes in detail, with reference to the drawings, the embodiments for implementing the information processing device, information processing method, and information processing program according to the present application (hereinafter referred to as "embodiments"). Note that these embodiments do not limit the information processing device, information processing method, and information processing program according to the present application. Furthermore, the same parts are denoted by the same reference numerals in each of the following embodiments, and redundant descriptions are omitted.

[0011] [1. Embodiments] The information processing realized by the information processing device of this embodiment will be explained using Figure 1. Figure 1 is a diagram showing an example of information processing according to the embodiment. In Figure 1, the information processing according to the embodiment is realized by the information processing device 10, which is an example of the information processing device according to the present application.

[0012] As shown in Figure 1, the information processing system 1 according to this embodiment includes an information processing device 10 and a user terminal 100. The information processing device 10 and the user terminal 100 are connected to each other via a network N (see, for example, Figure 2) by wire or wireless means so that they can communicate with each other. The network N is, for example, a WAN (Wide Area Network) such as the Internet. Note that the information processing system 1 shown in Figure 1 may include multiple information processing devices 10 and multiple user terminals 100.

[0013] The information processing device 10 shown in Figure 1 is an information processing device that performs information processing according to the embodiment, and can be realized by, for example, a server device or a cloud system. For example, the information processing device 10 receives video images from a user (in other words, a provider of video images) and provides a video editing service that adds captions to the received video images.

[0014] The information processing device 10 may also function as a web server providing a website related to the video editing service. Furthermore, the information processing device 10 may be a device that distributes information to be displayed in an application related to the video editing service installed on the user terminal 100. The information processing device 10 may also be a server that distributes the application data itself. Additionally, the information processing device 10 may function as a distribution device that distributes control information to the user terminal 100. Here, the control information is described, for example, using a scripting language such as JavaScript (registered trademark) or a stylesheet language such as CSS (Cascading Style Sheets). The application itself distributed from the information processing device 10 may also be considered as control information.

[0015] The user terminal 100 shown in Figure 1 is an information processing device used by a user. For example, the user terminal 100 can be a smartphone, a tablet, a notebook PC (Personal Computer), a desktop PC, a mobile phone, a PDA (Personal Digital Assistant), etc. In the example shown in Figure 1, the user terminal 100 is a smartphone used by the user.

[0016] Furthermore, the user terminal 100 displays the information provided by the information processing device 10 using a web browser or application. When the user terminal 100 receives control information from the information processing device 10 or other devices to implement the information display process, it performs the display process according to the control information.

[0017] The information processing performed by the information processing device 10 will be explained below using Figure 1. In the following explanation, it will be assumed that the user terminal 100 is used by a user (user U1) identified by the user ID "UID#1". In the following explanation, the user terminal 100 may be treated as the same as user U1. That is, in the following, user U1 can be read as user terminal 100.

[0018] First, user U1 captures a video using user terminal 100 (step S1). In the example in Figure 1, user U1 visits store #1 located in area #1 and captures a video C1 that includes the target OB1 (store #1), target OB2 (user U1), and target OB3 (tapioca) held by user U1.

[0019] Subsequently, the information processing apparatus 10 acquires, from the user terminal 100 via the moving image editing service, explanatory information that describes the content of the moving image C1 together with the moving image C1 (step S2). For example, the information processing apparatus 10 acquires, as the explanatory information, text information indicating the description (caption) of the moving image C1 input by the user U1. Here, the description may include, for example, the purpose of shooting the moving image C1, information indicating the shooting target, the shooting target, and information about the user U1 (for example, attribute information).

[0020] In addition, when the moving image C1 is shot, the information processing apparatus 10 acquires, as the explanatory information, position information measured by the user terminal 100 using GPS (Global Positioning System) or the like (that is, the position information when the moving image C1 is shot). Further, the information processing apparatus 10 may acquire additional information based on the position information when the moving image C1 is shot.

[0021] Here, the additional information may include information about facilities and stores related to the position information acquired from a server device or the like not shown separately (for example, detailed information about products and menus described on the homepage of the store, evaluations of the store by evaluation sites, SNS (Social Networking Service), and word-of-mouth information from users posted on websites, but not limited to this). As another example of the additional information, event information that was held at the location indicated by the position information when the moving image C1 was shot, acquired from a server device or the like not shown separately (for example, information about the schedule and implementation content described on the homepage of the event, evaluations of the event by evaluation sites, SNS, and word-of-mouth information from users posted on websites, but not limited to this) may be included.

[0022] Furthermore, the information processing device 10 acquires text information indicating the audio contained in the video C1 as explanatory information. For example, the information processing device 10 acquires text information indicating the audio SD1 spoken by user U1 as explanatory information. Note that such text information may be acquired by the user terminal 100 from the audio SD1 using any speech recognition technology. Alternatively, such text information may be acquired by the information processing device 10 from the audio SD1 using any speech recognition technology.

[0023] Furthermore, the information processing device 10 obtains the analysis results obtained by analyzing the video image C1 using any image analysis technique as explanatory information. Here, the analysis results may include text information such as the facial expression of the subject being filmed (e.g., user U1), the actions of the subject being filmed, an object being held by the subject being filmed (e.g., subject OB3), and a string of characters contained in the video image C1 (e.g., the string of characters "Store #1" shown on the sign for Store #1). Note that such analysis results may be obtained by the user terminal 100 from the video image C1 using any image analysis technique. Alternatively, such analysis results may be obtained by the information processing device 10 from the video image C1 using any image analysis technique.

[0024] Furthermore, the information processing device 10 may determine the likelihood of acquiring additional information (described above) based on the "text information indicating the voice SD1" and the "analysis results of the video C1," and decide whether or not to use the additional information in the processing described later. For example, if the information processing device 10 acquires a tapioca shop, which is a specific store related to location information, as additional information, it will determine a value indicating the likelihood of the acquired information about the tapioca shop relating to the video, based on the "text information indicating the voice SD1" and the "analysis results of the video C1." The information processing device 10 may then decide to use the additional information in the processing described later if the value indicating the likelihood is above a predetermined threshold (for example, the value indicating the likelihood will be high if the audio in the video C1 includes the phrase "I have come to tapioca shop XX," if the text "tapioca" is analyzed as a result of the analysis of the video C1, or if actual tapioca is shown in the video C1).

[0025] Next, the information processing device 10 generates a caption to be added to the video C1 based on the descriptive information of the video C1 (step S3). For example, the information processing device 10 generates a caption to be added to the video C1 using the descriptive information of the video C1 and model #1, which has been trained to generate answers to input questions.

[0026] To give a specific example, the information processing device 10 generates a caption by inputting explanatory information about the video image C1, a rule that specifies the generation of a caption showing information about the subject being filmed in the video image C1, and an instruction sentence that instructs the model #1 to generate the caption based on the explanatory information of the video image C1 according to the rule. For example, the information processing device 10 generates a caption T1 "Tapioca at store #1 in area #1" showing information about the subject being filmed OB3, based on explanatory information such as location information when the video image C1 was filmed (for example, area #1), text information indicating the voice SD1 spoken by user U1, the string "store #1" indicated by the sign of store #1, and the text information indicating the voice SD1, such as "I came to drink tapioca today."

[0027] Furthermore, the information processing device 10 may generate a caption indicating information about the subject of the video by using the above-mentioned additional information in addition to the explanatory information of the video C1. For example, the information processing device 10 may generate a caption (not shown) that says, "Today I came to tapioca shop XX, which is popular for its mango tapioca juice," based on explanatory information such as the text information indicating the audio SD1, "Today I came to drink tapioca," and the additional information.

[0028] Furthermore, the information processing device 10 generates a caption by inputting explanatory information for the video C1, a rule that specifies the generation of a caption indicating the feelings of the subject being filmed in the video C1, and an instruction sentence that instructs the model #1 to generate the caption based on the explanatory information for the video C1 according to the rule. For example, the information processing device 10 generates a caption T2 "Looks delicious!" indicating the feelings of the subject being filmed OB2, based on the facial expression of the subject being filmed OB2 and explanatory information such as the text information indicating the voice SD1, "I came to drink tapioca today."

[0029] Furthermore, the information processing device 10 generates a caption by inputting explanatory information for the video image C1, a rule that specifies the generation of a caption indicating the situation of the subject being filmed in the video image C1, and an instruction sentence that instructs the model #1 to generate the caption based on the explanatory information for the video image C1 according to the rule. For example, the information processing device 10 generates a caption T3 "I tripped..." indicating the situation of the subject being filmed OB2, based on explanatory information indicating the action of the subject being filmed OB2 (for example, "tripping").

[0030] Furthermore, the information processing device 10 may generate a caption indicating the situation of the subject being filmed, such as "I came to store #1," based on the location information (for example, area #1) at the time the video C1 was filmed, and descriptive information such as the string "store #1" displayed on the sign of store #1.

[0031] Furthermore, the information processing device 10 may generate a caption indicating the situation of the subject being filmed, such as a caption that shows the atmosphere around the subject being filmed. For example, the information processing device 10 may generate a caption indicating the degree of crowding around the subject being filmed, depending on the number of people included in the video C1. For example, if the number of people included in the video C1 is above a predetermined threshold (i.e., it is crowded), the information processing device 10 may generate a caption indicating that it is crowded, such as "It's crowded," or a caption showing onomatopoeia such as "It's lively." Also, if the number of people included in the video C1 is below a predetermined threshold (i.e., it is empty), the information processing device 10 may generate a caption indicating that it is empty, such as "It's empty," or a caption showing onomatopoeia such as "It's deserted."

[0032] Furthermore, the information processing device 10 may generate a caption indicating the atmosphere around the subject being filmed, such as a caption indicating the degree of excitement around the subject being filmed. For example, if the volume of sound contained in the video C1 is above a predetermined threshold (i.e., it is exciting), the information processing device 10 may generate a caption indicating excitement such as "It's exciting" or a caption indicating onomatopoeia such as "Wow, wow." Also, if the volume of sound contained in the video C1 is below a predetermined threshold (i.e., it is not exciting), the information processing device 10 may generate a caption indicating that it is not exciting such as "It's a calm atmosphere" or a caption indicating onomatopoeia such as "Silence."

[0033] Furthermore, the information processing device 10 may generate a caption that takes into account information about user U1. For example, the information processing device 10 may generate a caption with a writing style that corresponds to user U1's attribute information (e.g., demographic attributes or psychographic attributes). For example, the information processing device 10 may generate a caption by inputting the explanatory information of the video C1, the attribute information of user U1, a rule that specifies the generation of a caption with a writing style that corresponds to user U1's attribute information, and an instruction sentence that instructs the model #1 to generate a caption based on the explanatory information of the video C1 according to the rule.

[0034] Furthermore, the information processing device 10 may generate captions that take into account the purpose for which user U1 filmed the video C1. For example, the information processing device 10 may generate captions in a style appropriate to the purpose for which user U1 filmed the video C1 (for example, a so-called "project video" filmed for an advertising project, or "private" footage of user U1's private life). To give a specific example, the information processing device 10 may generate captions by inputting explanatory information about the video C1, the purpose for which user U1 filmed the video C1, a rule that specifies the generation of captions in a style appropriate to that purpose, and an instruction sentence that instructs the device to generate captions based on the explanatory information about the video C1 according to that rule into model #1. For example, if the purpose for which user U1 filmed the video C1 was a "project video," the information processing device 10 will generate captions in polite language. Also, if the purpose for which user U1 filmed the video C1 was "private," the information processing device 10 may generate captions using informal language.

[0035] Furthermore, the information processing device 10 may receive requests (rules) from user U1 regarding the captions and generate captions in accordance with the received requests. For example, the information processing device 10 may generate captions by inputting explanatory information for the video image C1, requests from user U1, rules specifying the generation of captions in accordance with those requests, and instruction text instructing the model #1 to generate captions based on the explanatory information for the video image C1 in accordance with those rules. For example, user U1's requests may include "I want the name of store #1 to be displayed in the captions" or "I don't want specific location information to be displayed."

[0036] Next, the information processing device 10 presents the generated captions T1, T2, T3, ... to user U1 via user terminal 100 (step S4). For example, the information processing device 10 presents captions T1, T2, T3, ... as candidates for captions to be added to the video C1.

[0037] Here, in the example in Figure 1, let's assume that user U1 wishes to add a modified version of the text overlay T2 as a text overlay to the video C1. In such a case, the information processing device 10 receives modification information indicating the modifications to the text overlay T2 from the user terminal 100 via the video editing service (step S5). For example, the information processing device 10 receives the modification information for the text overlay T2 from user U1 in a conversational format, outputting a response corresponding to the message entered by user U1 using generation AI (Artificial Intelligence).

[0038] Next, the information processing device 10 modifies the ticker T2 based on the modification information (step S6). Here, suppose the modification information instructs that the font of the ticker T2 be changed to a font specified by user U1. In this case, the information processing device 10 generates a ticker T21 in which the font of the ticker T2 has been changed to the font specified by user U1.

[0039] The information processing device 10 may also modify the text overlay T2 by inputting the text overlay T2, modification information indicating the image of the modification, rules specifying that the text overlay T2 should be modified according to the modification information, and instruction text instructing that the distribution method be modified based on the modification information according to the rules, into model #1. Here, the modification information may be information indicating the image of the modification, such as "I want a fun font" or "I want the text overlay color to be a more eye-catching color."

[0040] Next, the information processing device 10 adds a caption T21 to the video C1 (step S7). For example, the information processing device 10 adds a caption T21 to the video C1 and generates the video SC1.

[0041] The information processing device 10 may be assigned to a pre-set position (for example, the lower area of ​​the moving image C1), or to a position specified by the user U1.

[0042] Furthermore, the information processing device 10 may add to the video image C1 multiple titles selected by the user U1 from among titles T1, T2, T3, ... or titles that are modified versions of titles T1, T2, T3, ....

[0043] Next, the information processing device 10 provides the video SC1 to the user terminal 100 via the video editing service (step S8). For example, the information processing device 10 provides the data of the video SC1 to the user terminal 100.

[0044] The information processing device 10 may also provide information indicating a caption (for example, text information) to the user terminal 100. The user U1 may then use video editing software or the like installed on the user terminal 100 to add the provided caption to the video C1.

[0045] As described above, the information processing device 10 according to the embodiment generates captions to be added to the video based on explanatory information describing the content of the video. This allows the information processing device 10 according to the embodiment to generate and add appropriate captions to the content of the video.

[0046] Furthermore, in recent years, the posting of videos by users on social media and other platforms has become widespread, and technologies for adding subtitles to such videos have been proposed. However, conventional technologies merely perform audio recognition on videos and automatically display audio subtitles, and have not disclosed how to consider the content (context) of the video and add subtitles that indicate the subject of the video, the subject's feelings, or the subject's situation.

[0047] Therefore, according to the information processing device 10 of this embodiment, it is possible to add captions that show the subject of the video, the subject's state of mind, the subject's condition, etc., taking into account the content of the video, thus saving the user the trouble of thinking up captions from scratch and improving convenience.

[0048] [2. Other processing examples] The above-described process is merely an example, and the information processing device 10 may perform various processes using various types of information. Examples of this are listed below.

[0049] [2-1. Regarding the display method of on-screen text] In the example shown in Figure 1, the information processing device 10 may generate captions in various display modes. For example, the information processing device 10 may use a color appropriate to the subject OB3, such as the caption T1 "Tapioca at store #1 in area #1," which indicates the subject OB3, (for example, a brown color that evokes the image of "tapioca"). The information processing device 10 may also use a high-brightness color for captions indicating positive emotions of the subject or situations that evoke positive emotions in the subject (for example, caption T2 "Looks delicious!"). Furthermore, the information processing device 10 may use a low-brightness color for captions indicating negative emotions of the subject or situations that evoke negative emotions in the subject (for example, caption T3 "I tripped..."). The information processing device 10 may also determine the color of the captions based on the emotions of the subject OB2, which are estimated from the voice SD1 or the subject OB2's facial expression.

[0050] Furthermore, the information processing device 10 may use a pop font for captions indicating positive emotions of the subject being filmed or situations that evoke positive emotions in the subject being filmed. Conversely, the information processing device 10 may use a dark font for captions indicating negative emotions of the subject being filmed or situations that evoke negative emotions in the subject being filmed. The information processing device 10 may also determine the font of the captions based on the emotions of the subject being filmed, which are estimated from the audio SD1 and the facial expressions of the subject being filmed OB2.

[0051] Furthermore, the information processing device 10 may make the size of the captions indicating positive emotions of the subject being filmed, or situations that evoke positive emotions in the subject being filmed, larger than usual (in other words, larger than captions that do not indicate positive emotions or situations that evoke positive emotions in the subject being filmed). Also, the information processing device 10 may make the size of the captions indicating negative emotions of the subject being filmed, or situations that evoke negative emotions in the subject being filmed, smaller than usual (in other words, smaller than captions that do not indicate negative emotions or situations that evoke negative emotions in the subject being filmed). The information processing device 10 may also determine the size of the captions based on the emotions of the subject being filmed, estimated from the audio SD1 and the facial expressions of the subject being filmed OB2.

[0052] Furthermore, the information processing device 10 may set movement (motion) on the caption depending on whether it indicates positive emotions of the subject being filmed or situations that evoke positive emotions of the subject being filmed, or negative emotions of the subject or situations that evoke negative emotions of the subject being filmed. The information processing device 10 may also set movement on the caption according to the situation (action) of the subject being filmed that the caption indicates. For example, the information processing device 10 may set movement (for example, tilting) that evokes the image of tripping for caption T3 "I tripped...".

[0053] [3. Configuration of the Information Processing Device] Next, the configuration of the information processing device 10 will be described using Figure 2. Figure 2 is a diagram showing an example of the configuration of the information processing device 10 according to the embodiment. As shown in Figure 2, the information processing device 10 has a communication unit 20, a storage unit 30, and a control unit 40.

[0054] (Regarding Communications Section 20) The communication unit 20 is implemented, for example, by a NIC (Network Interface Card). The communication unit 20 is connected to the network N by wire or wireless connection and performs information transmission and reception, for example, with the user terminal 100.

[0055] (Regarding memory unit 30) The storage unit 30 is implemented by, for example, semiconductor memory elements such as RAM (Random Access Memory) and flash memory, or by storage devices such as hard disks and optical discs. As shown in Figure 2, the storage unit 30 has a video information database 31 and a model database 32.

[0056] (Regarding the video information database 31) The video information database 31 stores various types of information related to videos (for example, descriptive information that explains the content of the video). Here, an example of the information stored in the video information database 31 will be explained using Figure 3. Figure 3 is a diagram showing an example of the video information database 31. In the example in Figure 3, the video information database 31 has items such as "User ID", "Attribute Information", "Video ID", "Video", "Description", and "Location Information".

[0057] "User ID" indicates identification information used to identify the user. "Attribute Information" indicates the user's demographic and psychographic attributes. "Video ID" indicates identification information used to identify videos provided by the user (for example, videos taken by the user). "Video" indicates the video provided by the user. "Description" indicates the description entered by the user about the video. "Location Information" indicates the location information at the time the video was taken (for example, the location information determined by user terminal 100 that took the video).

[0058] In other words, Figure 3 shows an example where the user attribute information identified by the user ID "UID#1" is "Attribute Information #1", the video provided by the user "Video #1" is identified by the video ID "CID#1", the description is "Description #1", and the location information is "Location Information #1".

[0059] The information stored in the video information database 31 is not limited to the above. For example, the video information database 31 may store information indicating the purpose for which the video C1 was filmed, information indicating the subject of filming, and text information indicating the sound contained in the video. In addition, the video information database 31 may store text information indicating the facial expression of the subject of filming, the actions of the subject of filming, the object being held by the subject of filming, and strings of characters contained in the video.

[0060] (Regarding Model Database 32) The model database 32 stores models that have been trained to generate answers to input questions.

[0061] (Regarding the control unit 40) The control unit 40 is a controller, and is realized, for example, by a CPU (Central Processing Unit) or MPU (Micro Processing Unit) executing various programs stored in the memory device inside the information processing device 10 using RAM as a working area. Alternatively, the control unit 40 is a controller, and is realized, for example, by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array). As shown in Figure 2, the control unit 40 according to this embodiment has an acquisition unit 41, a generation unit 42, a reception unit 43, a modification unit 44, and an assignment unit 45, and realizes or executes the information processing functions and operations described below.

[0062] (Regarding acquisition section 41) The acquisition unit 41 acquires descriptive information that explains the content of the video. For example, in the example in Figure 1, the acquisition unit 41 acquires the video C1 along with descriptive information that explains the content of the video C1 from the user terminal 100 via the video editing service and stores it in the storage unit 30 (for example, the video information database 31).

[0063] (Regarding the generation unit 42) The generation unit 42 generates captions to be added to the video based on the explanatory information acquired by the acquisition unit 41. For example, in the example in Figure 1, the generation unit 42 refers to the storage unit 30 (for example, the video information database 31) and generates captions to be added to the video C1 based on the explanatory information of the video C1.

[0064] Furthermore, the generation unit 42 may generate a caption indicating the subject of the video. For example, in the example in Figure 1, the generation unit 42 generates a caption T1 "Tapioca at store #1 in area #1" indicating information about the subject OB3, based on location information when the video C1 was taken, text information indicating the voice SD1 spoken by user U1, the string "Store #1" indicated by the sign for store #1, and explanatory information such as the text information indicating the voice SD1 "I came to drink tapioca today."

[0065] Furthermore, the generation unit 42 may generate a caption indicating the feelings of the subject being filmed in the video. For example, in the example in Figure 1, the generation unit 42 generates a caption T2 "Looks delicious!" indicating the feelings of the subject OB2, based on the facial expression of the subject OB2 and explanatory information such as the text information indicating the voice SD1, "I came to drink tapioca today."

[0066] Furthermore, the generation unit 42 may generate a caption indicating the status of the subject being filmed in the video. For example, in the example in Figure 1, the generation unit 42 generates a caption T3 "I tripped..." indicating the status of the subject OB2, based on explanatory information indicating the operation of the subject OB2.

[0067] Furthermore, the generation unit 42 may generate a caption with a display mode corresponding to the subject of the moving image. For example, in the example in Figure 1, the generation unit 42 displays the caption T1 "Tapioca at store #1 in area #1" which indicates the subject OB3, in a color corresponding to the subject OB3.

[0068] Furthermore, the generation unit 42 may generate captions that display according to the emotional state of the subject being filmed in the video. For example, in the example in Figure 1, the generation unit 42 uses a high-brightness color for captions indicating positive emotions of the subject being filmed. The generation unit 42 also uses a low-brightness color for captions indicating negative emotions of the subject being filmed. Furthermore, the generation unit 42 uses a pop font for captions indicating positive emotions of the subject being filmed. Furthermore, the generation unit 42 uses a dark-looking font for captions indicating negative emotions of the subject being filmed. Furthermore, the generation unit 42 makes the size of captions indicating positive emotions of the subject being filmed larger than usual. Furthermore, the generation unit 42 makes the size of captions indicating negative emotions of the subject being filmed smaller than usual. In addition, the generation unit 42 sets movement for the captions depending on whether the caption indicates positive or negative emotions of the subject being filmed.

[0069] Furthermore, the generation unit 42 may generate captions in a display manner that corresponds to the situation of the subject being filmed in the video. For example, in the example in Figure 1, the generation unit 42 uses a high-brightness color for captions indicating situations that evoke positive emotions towards the subject. The generation unit 42 also uses a low-brightness color for captions indicating situations that evoke negative emotions towards the subject. Furthermore, the generation unit 42 uses a pop font for captions indicating situations that evoke positive emotions towards the subject. Furthermore, the generation unit 42 uses a dark-looking font for captions indicating situations that evoke negative emotions towards the subject. Furthermore, the generation unit 42 makes the size of captions indicating situations that evoke positive emotions towards the subject larger than usual. Furthermore, the generation unit 42 makes the size of captions indicating situations that evoke negative emotions towards the subject smaller than usual. Furthermore, the generation unit 42 sets movement for the text depending on whether the text indicates a situation that evokes positive or negative emotions towards the subject being filmed. In addition, the generation unit 42 sets movement for the text T3 "I tripped..." that represents the image of tripping.

[0070] Furthermore, the generation unit 42 may generate a caption by inputting explanatory information and an instruction sentence that instructs a model trained to output an answer to an input question to output a caption for the video based on the explanatory information. For example, in the example in Figure 1, the generation unit 42 refers to the storage unit 30 (for example, the video information database 31 or the model database 32) and uses the explanatory information of the video C1 and model #1, which has been trained to generate an answer to an input question, to generate a caption to be attached to the video C1.

[0071] Furthermore, the generation unit 42 may generate a caption by inputting explanatory information, rules specified by the provider of the video, and instruction text that instructs the model to output a caption for the video based on the explanatory information in accordance with the rules. For example, in the example in Figure 1, the generation unit 42 may generate a caption by inputting explanatory information for the video C1, a request from user U1, rules that specify the generation of a caption corresponding to that request, and instruction text that instructs the model to generate a caption based on the explanatory information for the video C1 in accordance with the rules.

[0072] Furthermore, the generation unit 42 may generate multiple captions to be added to the video. For example, in the example in Figure 1, the generation unit 42 generates captions T1, T2, T3, ... to be added to the video C1.

[0073] Furthermore, the generation unit 42 may generate background music (BGM) to be added to the video based on the acquired explanatory information and / or additional information. For example, if the generation unit 42 has acquired information about the text or audio of "tapioca" and information about a tapioca shop as additional information, it may generate and add BGM that evokes the image of tapioca (for example, a poppy tune or a tune that evokes the image of a tropical country).

[0074] Furthermore, the generation unit that generates background music (BGM) and the generation unit 42 that generates the captions described above do not have to be physically identical. For example, the BGM generated here may be generated by a server device (not shown) different from the information processing device 10, and the acquisition of BGM generated by that different server device by the information processing device 10 (generation unit 42) is also included in the generation of the BGM. In addition, the BGM generated here does not necessarily have to be mechanically generated using AI or the like; for example, it may be royalty-free sound sources that can be obtained from the internet.

[0075] (Regarding Reception Desk 43) The reception unit 43 receives correction information indicating the content of the corrections to the text overlay from the video provider. For example, in the example in Figure 1, the reception unit 43 receives correction information indicating the content of the corrections to the text overlay T2 from the user terminal 100 via the video editing service.

[0076] The reception unit 43 may also accept correction information in a conversational format with the provider. For example, in the example in Figure 1, the reception unit 43 accepts correction information for the text overlay T2 from user U1 in a conversational format, outputting a response corresponding to the message entered by user U1 using a generation AI.

[0077] (Regarding amendment 44) The modification unit 44 modifies the text overlay based on the modification information received by the reception unit 43. For example, in the example in Figure 1, the modification unit 44 modifies the text overlay T2 based on the modification information received from the user terminal 100.

[0078] Furthermore, the modification unit 44 may modify the on-screen text by inputting modification information, the on-screen text, and an instruction sentence instructing the model, which has been trained to output an answer to the input question, to modify the on-screen text based on the modification information. For example, in the example in Figure 1, the modification unit 44 modifies the on-screen text T2 by referring to the storage unit 30 (e.g., the model database 32) and inputting the on-screen text T2, modification information showing an image of the modification, a rule specifying that the on-screen text T2 should be modified according to the modification information, and an instruction sentence instructing the model #1 to modify the delivery method based on the modification information according to the rule.

[0079] (Regarding the attachment section 45) The assignment unit 45 assigns the caption generated by the generation unit 42 to the video. For example, in the example in Figure 1, the assignment unit 45 assigns the caption T21 to the video C1 and generates the video SC1.

[0080] Furthermore, the assignment unit 45 may assign to the video video a caption selected by the video provider from among multiple captions. For example, in the example in Figure 1, the assignment unit 45 assigns to the video video C1 multiple captions selected by user U1 from among captions T1, T2, T3, ...

[0081] Furthermore, the assignment unit 45 may assign the text overlay that has been modified by the modification unit 44 to the video. For example, in the example in Figure 1, the assignment unit 45 assigns the text overlay T21 that has been modified by the modification unit 44 to the video C1.

[0082] [4. Information Processing Flow] The information processing procedure of the information processing device 10 according to the embodiment will be explained using Figure 4. Figure 4 is a flowchart showing an example of the information processing procedure according to the embodiment.

[0083] As shown in Figure 4, the information processing device 10 determines whether or not it has acquired explanatory information describing the content of the video (step S101). If explanatory information has not been acquired (step S101; No), the information processing device 10 waits until it acquires the explanatory information.

[0084] On the other hand, if explanatory information is obtained (step S101; Yes), the information processing device 10 generates a caption to be added to the video based on the explanatory information (step S102). Subsequently, the information processing device 10 adds the caption to the video (step S103) and terminates the process.

[0085] [5. Variations] The above-described embodiment is merely an example, and various modifications and applications are possible.

[0086] [5-1. Regarding the processing method] Of the processes described in the above embodiments, all or part of the processes described as being performed automatically can be performed manually, and conversely, all or part of the processes described as being performed manually can be performed automatically by known methods. In addition, the processing procedures, specific names, and information including various data and parameters shown in the above text and drawings can be arbitrarily changed unless otherwise specified. For example, the various information shown in each figure is not limited to the information shown.

[0087] Furthermore, the components of each illustrated device are functionally conceptual and do not necessarily need to be physically configured as shown. In other words, the specific forms of distribution and integration of each device are not limited to those shown, and all or part of them can be functionally or physically distributed and integrated in any unit according to various loads and usage conditions.

[0088] Furthermore, the embodiments described above can be combined as appropriate, provided that the processing content is not contradictory.

[0089] [6. Effects] As described above, the information processing device 10 according to the embodiment includes an acquisition unit 41, a generation unit 42, a reception unit 43, a modification unit 44, and an assignment unit 45. The acquisition unit 41 acquires explanatory information that describes the content of a video. The generation unit 42 generates captions to be added to the video based on the explanatory information acquired by the acquisition unit 41. The generation unit 42 also generates captions by inputting explanatory information and instruction sentences that instruct a model trained to output answers to input questions to output captions for the video based on the explanatory information. The generation unit 42 also generates captions by inputting explanatory information, rules specified by the video provider, and instruction sentences that instruct the model to output captions for the video based on the explanatory information according to the rules. The reception unit 43 receives modification information from the video provider indicating the content to be modified for the captions. The reception unit 43 also receives modification information in a conversational format with the provider. The modification unit 44 modifies the captions based on the modification information received by the reception unit 43. Furthermore, the modification unit 44 modifies the on-screen text by inputting modification information, the on-screen text, and an instruction sentence instructing the model, which has been trained to output an answer to the input question, to modify the on-screen text based on the modification information. The application unit 45 applies the on-screen text generated by the generation unit 42 to the video. The application unit 45 also applies the on-screen text selected by the video provider from among multiple on-screen texts to the video. The application unit 45 also applies the on-screen text that has been modified by the modification unit 44 to the video.

[0090] As a result, the information processing device 10 according to the embodiment can generate captions to be added to the video based on explanatory information describing the content of the video, and can generate and add appropriate captions to the content of the video.

[0091] Furthermore, in the information processing device 10 according to the embodiment, for example, the generation unit 42 generates a caption indicating the subject of the video. The generation unit 42 also generates a caption indicating the state of mind of the subject of the video. The generation unit 42 also generates a caption indicating the situation of the subject of the video. The generation unit 42 also generates a caption in a display manner corresponding to the subject of the video. The generation unit 42 also generates a caption in a display manner corresponding to the state of mind of the subject of the video. The generation unit 42 also generates a caption in a display manner corresponding to the situation of the subject of the video. Furthermore, the generation unit 42 generates a caption by inputting explanatory information and an instruction sentence instructing a model trained to output answers to input questions to output captions for the video based on the explanatory information. Furthermore, the generation unit 42 generates a caption by inputting explanatory information, rules specified by the video provider, and an instruction sentence instructing the model to output captions for the video based on the explanatory information in accordance with the rules.

[0092] As a result, the information processing device 10 according to the embodiment can generate various types of captions and add them to moving images, thereby improving convenience.

[0093] [7. Hardware Configuration] Furthermore, the information processing device 10 according to each embodiment described above can be implemented by a computer 1000 having a configuration such as that shown in Figure 5. The following explanation will use the information processing device 10 as an example. Figure 5 is a hardware configuration diagram showing an example of a computer that implements the functions of the information processing device 10. The computer 1000 has a CPU 1100, ROM 1200, RAM 1300, HDD 1400, communication interface (I / F) 1500, input / output interface (I / F) 1600, and media interface (I / F) 1700.

[0094] The CPU 1100 operates based on programs stored in the ROM 1200 or HDD 1400, and controls various parts. The ROM 1200 stores boot programs executed by the CPU 1100 when the computer 1000 starts up, as well as programs that depend on the computer 1000's hardware.

[0095] The HDD 1400 stores programs executed by the CPU 1100, as well as data used by such programs. The communication interface 1500 receives data from other devices via the communication network 500 (corresponding to network N in this embodiment) and sends it to the CPU 1100, and also transmits data generated by the CPU 1100 to other devices via the communication network 500.

[0096] The CPU 1100 controls output devices such as displays and printers, and input devices such as keyboards and mice, via the input / output interface 1600. The CPU 1100 acquires data from input devices via the input / output interface 1600. The CPU 1100 also outputs the data it generates to output devices via the input / output interface 1600.

[0097] The media interface 1700 reads a program or data stored in the recording medium 1800 and provides it to the CPU 1100 via the RAM 1300. The CPU 1100 loads the program from the recording medium 1800 onto the RAM 1300 via the media interface 1700 and executes the loaded program. The recording medium 1800 is, for example, an optical recording medium such as a DVD (Digital Versatile Disc) or PD (Phase Change Rewritable Disk), a magneto-optical recording medium such as an MO (Magneto-Optical disk), a tape medium, a magnetic recording medium, or a semiconductor memory.

[0098] For example, when computer 1000 functions as information processing device 10, the CPU 1100 of computer 1000 realizes the functions of control unit 40 by executing programs loaded on RAM 1300. The HDD 1400 stores the data in the storage device of information processing device 10. The CPU 1100 of computer 1000 reads and executes these programs from the recording medium 1800, but as another example, these programs may be obtained from other devices via a predetermined communication network.

[0099] [8. Other] Although some embodiments of the present invention have been described in detail above with reference to the drawings, these are illustrative examples, and the present invention can be implemented in various other forms with modifications and improvements based on the knowledge of those skilled in the art, starting with the embodiments described in the disclosure section of the invention.

[0100] Furthermore, the configuration of the aforementioned information processing device 10 can be flexibly changed, for example, by calling external platforms, etc., via APIs (Application Programming Interfaces) or network computing, depending on the function.

[0101] Furthermore, the term "part" in the claims can be replaced with "means," "circuit," etc. For example, "reception part" can be replaced with "reception means" or "reception circuit." [Explanation of Symbols]

[0102] 1. Information Processing System 10 Information Processing Devices 20 Communications Department 30 Storage section 31. Video Information Database 32 Model Databases 40 Control Unit 41 Acquisition Department 42 Generation part 43 Reception Department 44 Correction section 45 Granting section 100 User Terminals

Claims

1. An acquisition unit that acquires explanatory information describing the content of a video, A generation unit generates a caption to be added to the video image based on the explanatory information acquired by the acquisition unit, A labeling unit that applies the text generated by the generation unit to the video image. An information processing device characterized by having the following features.

2. The generating unit is The text overlay indicating the subject of the video is generated. The information processing apparatus according to feature 1.

3. The generating unit is The video generates the caption that shows the state of mind of the subject being filmed. The information processing apparatus according to feature 1.

4. The generating unit is The text overlay is generated to show the status of the subject being captured in the video. The information processing apparatus according to feature 1.

5. The generating unit is The system generates the caption in a display format corresponding to the subject of the video. The information processing apparatus according to feature 1.

6. The generating unit is The video generates a caption that displays according to the emotional state of the subject being filmed. The information processing apparatus according to feature 1.

7. The generating unit is The system generates the caption in a display mode that corresponds to the situation of the subject being captured in the video. The information processing apparatus according to feature 1.

8. The generating unit is The model, which has been trained to output answers to input questions, is given the explanatory information and an instruction sentence that instructs the model to output captions for the video based on the explanatory information, thereby generating the captions. The information processing apparatus according to feature 1.

9. The generating unit is The model is generated by inputting the explanatory information, the rules specified by the provider of the video, and an instruction document that instructs the model to output video captions based on the explanatory information according to the rules. The information processing apparatus according to feature 8.

10. The generating unit is Multiple captions are generated to be added to the aforementioned video image. The aforementioned attachment unit is, From among the multiple captions, the caption selected by the provider of the video is added to the video. The information processing apparatus according to feature 1.

11. A receiving unit that receives correction information indicating the content of the corrections to the aforementioned text overlay from the provider of the video image, Based on the correction information received by the reception unit, the correction unit corrects the on-screen text. It further possesses, The aforementioned attachment unit is, The text overlay, which has been modified by the aforementioned modification unit, is added to the video image. The information processing apparatus according to feature 1.

12. The aforementioned modification section is, The model, which has been trained to output answers to input questions, is modified by inputting the correction information, the caption, and an instruction sentence that instructs the model to modify the caption based on the correction information. The information processing apparatus according to feature 11.

13. The aforementioned reception unit is The correction information is received in the form of a conversation with the aforementioned provider. The information processing apparatus according to feature 11.

14. A method of information processing performed by a computer, The acquisition process involves obtaining explanatory information that describes the content of the video, A generation step, based on the explanatory information acquired in the acquisition step, generates a caption to be added to the video image, A process of applying the text generated by the generation process to the video image. An information processing method characterized by including

15. Procedure for obtaining descriptive information that explains the content of a video, A generation procedure for generating captions to be added to the video image based on the explanatory information obtained by the acquisition procedure, A procedure for adding the caption generated by the above generation procedure to the video image, and An information processing program that causes a computer to execute something.

Citation Information

Patent Citations

  • Image processing device, image processing method, and program

    WO2019230225A1