Method, computer device, and computer program for triggering generation of video content on basis of conversation content
By analyzing conversation content and user-held video context, the method generates contextually relevant video content, addressing the limitations of existing instant messenger applications in creating meaningful video content.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- LINE PLUS
- Filing Date
- 2025-10-16
- Publication Date
- 2026-05-21
AI Technical Summary
Existing instant messenger applications lack the ability to generate meaningful video content based on conversation context and user-held photos, relying on simple thematic rules that do not consider the actual user intent.
A method and device that analyze conversation content using deep learning and text embeddings to identify key keywords and topics, cluster user-held videos, and generate video content by selecting similar videos based on both conversation context and user-held photo context, using similarity thresholds and additional triggers.
Generates timely and relevant video content that aligns with user intent, enhancing the usability of user-held photos by creating contextually meaningful videos.
Smart Images

Figure KR2025016422_21052026_PF_FP_ABST
Abstract
Description
Method for triggering video content generation based on conversation content, computer device, and computer program
[0001] The following description concerns the technology for creating video content.
[0002] Instant messengers, a common communication tool, are software capable of sending and receiving messages or data in real time, enabling users to register conversation partners as friends and exchange messages in real time with those in their contact list.
[0003] These messenger functions are becoming commonplace not only on PCs but also in the mobile environment of mobile communication devices.
[0004] For example, Korean Published Patent No. 10-2002-0074304 (published on September 30, 2002) discloses a mobile messenger service system and method for a mobile terminal using a wireless communication network that enables the provision of messenger services between mobile messengers installed on the mobile terminal.
[0005] The use of instant messengers is becoming more popular, and the features provided through instant messengers are becoming increasingly diverse.
[0006] You can create more meaningful video content based on messenger conversations.
[0007] Based on the content of the messenger conversation, a topic can be derived, and the creation of video content related to that topic can be triggered.
[0008] Users can classify and manage their own photos used to create video content by topic.
[0009] A method for triggering video content generation in a computer device comprising at least one processor, the method comprising: a step of analyzing the conversation content of a messenger by the at least one processor; and a step of triggering the generation of video content using the searched video when at least one video among the videos held by the user is searched by the at least one processor and has a similarity value of at least a threshold value with respect to the analysis result of the conversation content.
[0010] According to one aspect, the analyzing step may include: a step of determining a conversation section to be analyzed; and a step of analyzing a conversation included in the conversation section.
[0011] According to another aspect, the above-mentioned analysis step may include the step of extracting key keywords that serve as the topic of conversation from the above-mentioned conversation content.
[0012] According to another aspect, the above-mentioned analyzing step may include a step of generating a summary sentence by summarizing the conversation content through a deep learning model.
[0013] According to another aspect, the analysis step may include a step of converting the conversation content into text embeddings through a text encoder.
[0014] According to another aspect, the analyzing step may include a step of generating a summary sentence by summarizing the conversation content through a deep learning model; and a step of converting the summary sentence into a text embedding through a text encoder.
[0015] According to another aspect, the video content generation triggering method may further include the step of analyzing the user-held video and clustering it by topic by the at least one processor.
[0016] According to another aspect, the clustering step may generate a tag for each of the user-held images and cluster the user-held images by tag.
[0017] According to another aspect, the clustering step may extract features for each of the user-held images and cluster the user-held images based on the similarity between the features.
[0018] According to another aspect, the triggering step may include: a step of calculating the similarity between information obtained through the analysis of the conversation content and information obtained through the analysis of the user-held video; and a step of selecting a video in which the similarity is greater than or equal to a threshold value as a video associated with the conversation content.
[0019] According to another aspect, the triggering step may trigger the generation of the image content when an image having a similarity greater than or equal to the threshold value is found, and a predefined keyword or situation is detected.
[0020] A computer program stored on a computer-readable recording medium is provided to execute the above-mentioned video content generation triggering method on the computer device.
[0021] The present invention provides a computer device comprising at least one processor implemented to execute a readable command on a computer device, wherein the at least one processor processes the process of analyzing the conversation content of a messenger; and the process of triggering the creation of video content using the found video when at least one video among the videos held by the user has a similarity value greater than or equal to a reference value with respect to the analysis result of the conversation content.
[0022] FIG. 1 is a drawing illustrating an example of a network environment according to an embodiment of the present invention.
[0023] FIG. 2 is a block diagram illustrating an example of a computer device according to an embodiment of the present invention.
[0024] FIG. 3 is a flowchart illustrating an example of a method that a computer device according to an embodiment of the present invention can perform.
[0025] FIGS. 4 and 5 illustrate an example of an image management process in an embodiment of the present invention.
[0026] FIGS. 6 to 8 illustrate an example of a conversation content analysis process in an embodiment of the present invention.
[0027] FIG. 9 illustrates an example of a process for generating conversation-associated video content in an embodiment of the present invention.
[0028] FIGS. 10 to 12 illustrate examples of conversation-associated video content recommendations in an embodiment of the present invention.
[0029] Hereinafter, embodiments of the present invention will be described in detail with reference to the attached drawings.
[0030] Embodiments of the present invention relate to a technology for generating video content.
[0031] Embodiments including those specifically disclosed in this specification can generate more meaningful video content by triggering the generation of video content by taking into account both the context of the conversation in which the user is participating and the context of the photos held by the user.
[0032] A video content generation triggering device according to embodiments of the present invention may be implemented by at least one computer device, and a video content generation triggering method according to embodiments of the present invention may be performed through at least one computer device included in the video content generation triggering device. At this time, a computer program according to an embodiment of the present invention may be installed and run on the computer device, and the computer device may perform a video content generation triggering method according to embodiments of the present invention under the control of the run computer program. The above-described computer program may be stored on a computer-readable recording medium to be combined with the computer device to execute the video content generation triggering method on the computer.
[0033] FIG. 1 is a diagram illustrating an example of a network environment according to an embodiment of the present invention. The network environment of FIG. 1 illustrates an example including a plurality of electronic devices (110, 120, 130, 140), a plurality of servers (150, 160), and a network (170). FIG. 1 is an example for explaining the invention, and the number of electronic devices or servers is not limited to that shown in FIG. 1. Furthermore, the network environment of FIG. 1 is merely an example of one of the environments applicable to the present embodiments, and the environments applicable to the present embodiments are not limited to the network environment of FIG. 1.
[0034] Multiple electronic devices (110, 120, 130, 140) may be fixed terminals or mobile terminals implemented as computer devices. Examples of multiple electronic devices (110, 120, 130, 140) include smartphones, mobile phones, navigation systems, computers, laptops, digital broadcasting terminals, PDAs (Personal Digital Assistants), PMPs (Portable Multimedia Players), tablet PCs, etc. For example, FIG. 1 shows the shape of a smartphone as an example of an electronic device (110), but in embodiments of the present invention, the electronic device (110) may substantially refer to one of various physical computer devices capable of communicating with other electronic devices (120, 130, 140) and / or servers (150, 160) via a network (170) using a wireless or wired communication method.
[0035] The communication method is not limited and may include not only communication methods utilizing communication networks (e.g., mobile communication networks, wired internet, wireless internet, broadcasting networks) that the network (170) may include, but also short-range wireless communication between devices. For example, the network (170) may include any one or more networks such as a PAN (personal area network), LAN (local area network), CAN (campus area network), MAN (metropolitan area network), WAN (wide area network), BBN (broadband network), and the Internet. Additionally, the network (170) may include any one or more network topologies such as a bus network, a star network, a ring network, a mesh network, a star-bus network, a tree or hierarchical network, but is not limited thereto.
[0036] Each of the servers (150, 160) may be implemented as a computer device or multiple computer devices that communicate with multiple electronic devices (110, 120, 130, 140) through a network (170) to provide commands, code, files, content, services, etc. For example, the server (150) may be a system that provides services (e.g., messenger services, etc.) to multiple electronic devices (110, 120, 130, 140) connected through the network (170).
[0037] FIG. 2 is a block diagram illustrating an example of a computer device according to an embodiment of the present invention. Each of the plurality of electronic devices (110, 120, 130, 140) or servers (150, 160) described above can be implemented by the computer device (200) illustrated in FIG. 2.
[0038] As illustrated in FIG. 2, such a computer device (200) may include memory (210), a processor (220), a communication interface (230), and an input / output interface (240). The memory (210) is a computer-readable recording medium and may include a non-perishable mass storage device such as RAM (random access memory), ROM (read only memory), and a disk drive. Here, a non-perishable mass storage device such as a ROM and a disk drive may be included in the computer device (200) as a separate permanent storage device distinct from the memory (210). Additionally, an operating system and at least one program code may be stored in the memory (210). These software components may be loaded into the memory (210) from a computer-readable recording medium separate from the memory (210). This separate computer-readable recording medium may include a computer-readable recording medium such as a floppy drive, a disk, a tape, a DVD / CD-ROM drive, or a memory card. In another embodiment, software components may be loaded into memory (210) via a communication interface (230) rather than a computer-readable recording medium. For example, software components may be loaded into memory (210) of a computer device (200) based on a computer program installed by files received through a network (170).
[0039] The processor (220) may be configured to process instructions of a computer program by performing basic arithmetic, logic, and input / output operations. Instructions may be provided to the processor (220) via memory (210) or a communication interface (230). For example, the processor (220) may be configured to execute instructions received according to program code stored in a recording device such as memory (210).
[0040] The communication interface (230) may provide a function for the computer device (200) to communicate with other devices (e.g., storage devices described above) through the network (170). For example, requests, commands, data, files, etc. generated by the processor (220) of the computer device (200) according to program code stored in a recording device such as memory (210) may be transmitted to other devices through the network (170) under the control of the communication interface (230). Conversely, signals, commands, data, files, etc. from other devices may be received by the computer device (200) through the communication interface (230) of the computer device (200) via the network (170). Signals, commands, data, etc. received through the communication interface (230) may be transmitted to the processor (220) or memory (210), and files, etc. may be stored in a storage medium (the permanent storage device described above) that the computer device (200) may further include.
[0041] The input / output interface (240) may be a means for interfacing with an input / output device (250). For example, the input device may include a device such as a microphone, keyboard, or mouse, and the output device may include a device such as a display or speaker. As another example, the input / output interface (240) may be a means for interfacing with a device in which the functions for input and output are integrated into one, such as a touchscreen. The input / output device (250) may be composed of a computer device (200) and a single device.
[0042] Additionally, in other embodiments, the computer device (200) may include fewer or more components than the components of FIG. 2. However, it is not necessary to clearly illustrate most of the prior art components. For example, the computer device (200) may be implemented to include at least some of the input / output devices (250) described above, or may include other components such as a transceiver, a database, etc.
[0043] Hereinafter, specific embodiments of a method and device for generating video content based on conversation content will be described.
[0044] General photo applications include not only the function of automatically organizing all photos and videos into a timeline in one place, but also a memory feature that creates and recommends a single album for photos and videos with common information.
[0045] These recommended albums utilize various information such as the date, location, people, and season of the images (photos and videos), and are primarily generated according to predetermined rules. Recommended albums are created with titles based on specific dates (e.g., today one year ago, today two years ago, etc.) or general themes (e.g., autumn, cityscapes, etc.).
[0046] The embodiments may provide a method for triggering the generation of video content related to conversation content based on the conversation content of a messenger as one of the functions provided through a messenger platform, and a method for managing video data for this purpose.
[0047] In this specification, the term "image" may encompass photos or videos held by a user via a device or cloud, and the image content generated from such images may refer to content created such that multiple images are automatically played in a certain format (e.g., a slideshow).
[0048] The computer device (200) according to the present embodiment can provide a content recommendation service to a client by accessing a dedicated application installed on the client or a web / mobile site related to the computer device (200). The computer device (200) may be configured with a video content creation triggering device implemented by a computer. For example, the video content creation triggering device may be implemented in the form of a program that operates independently, or configured in the form of an in-app of a specific application so that it can operate on said specific application.
[0049] The processor (220) of the computer device (200) may be implemented as a component for performing the following video content generation triggering method. Depending on the embodiment, the components of the processor (220) may be optionally included in or excluded from the processor (220). Additionally, depending on the embodiment, the components of the processor (220) may be separated or merged to represent the function of the processor (220).
[0050] These processors (220) and components of the processor (220) can control a computer device (200) to perform steps included in the following video content generation triggering method. For example, the processor (220) and components of the processor (220) may be implemented to execute instructions according to the code of an operating system included in memory (210) and the code of at least one program.
[0051] Here, the components of the processor (220) may be representations of different functions performed by the processor (220) according to instructions provided by program code stored in the computer device (200).
[0052] The processor (220) can read necessary instructions from memory (210) in which instructions related to the control of the computer device (200) are loaded. In this case, the read instructions may include instructions for controlling the processor (220) to execute the steps to be described later.
[0053] The steps included in the video content generation triggering method described later may be performed in a different order than the one described, and some of the steps may be omitted or additional processes may be included.
[0054] The steps included in the video content generation triggering method may be performed on a client, and depending on the embodiment, at least some of the steps may also be performed on a server (150).
[0055] FIG. 3 is a flowchart illustrating an example of a method that a computer device according to an embodiment of the present invention can perform.
[0056] Referring to FIG. 3, in step (S310), the processor (220) can analyze the videos owned by the user and classify them based on the analysis results. In order to generate related video content based on the conversation content of the messenger, it is necessary to organize the videos on the user's device (or cloud) well in advance. The processor (220) can classify the videos owned by the user by topic in advance and build a cluster DB to facilitate the generation of video content. For example, the processor (220) can analyze the videos owned by the user and generate explicit meta-information from the video analysis results to perform clustering. As another example, the processor (220) can also perform clustering using similarity between videos without explicit information by extracting features (image embedding vectors) of the videos owned by the user.
[0057] In step (S320), the processor (220) can analyze the conversation content of the messenger and search for videos related to the conversation content based on the analysis results. In other words, the processor (220) can find videos related to the conversation content in the cluster DB for videos owned by the user. At this time, the processor (220) can first extract the conversation segment to be analyzed and then analyze the topic of the conversation segment. The result of the conversation content analysis may include at least one topic and may have the form of keywords, summary sentences, features, etc. Subsequently, the processor (220) can select a video having information similar to the result of the conversation content analysis from the videos in the cluster DB.
[0058] In step (S330), the processor (220) can perform video content generation using the video retrieved in step (S320). The processor (220) can trigger video content generation when at least one video is retrieved that has a similarity to the conversation content analysis result that is greater than or equal to a threshold value. According to an embodiment, in addition to the condition of similarity to the conversation content analysis result, video content generation can be initiated by detecting a situation where there is a request related to the video from a user or counterpart participating in the conversation, for example, when predefined keywords such as 'photo' or 'share' are entered into a conversation message or when an album access menu is executed. The video content can be generated to be automatically played in a certain format (e.g., a slideshow) composed of the video retrieved from the conversation content analysis result.
[0059] FIGS. 4 and 5 illustrate an example of an image management process in an embodiment of the present invention.
[0060] For example, the processor (220) can perform clustering by generating meta-information about the user's videos.
[0061] The processor (220) can generate at least one tag that can effectively represent the video using machine learning techniques, along with basic metadata such as the shooting date, shooting time, and shooting location for the videos owned by the user. Depending on the embodiment, it is also possible to use an annotation directly recorded by the user on the video as a tag.
[0062] The processor (220) can cluster user-held images by tag as illustrated in FIGS. 4 and FIGS. 5.
[0063] As another example, the processor (220) can perform clustering by extracting features from user-owned images. Features that can represent images well can be generated using deep learning or hand-crafted features.
[0064] Recently, feature extractor pairs capable of matching vision information and text information using deep learning-based VLMs (vision-language models) are widely used. In other words, given a specific image and text describing that image, the image and text can be converted into features through an image encoder and a text encoder, respectively. Since the similarity between these feature pairs (image features and text features) is very high, it is possible to find the text describing the image through the image or to find the image through the text.
[0065] When clustering user-owned images through feature extraction, feature information of the images can be stored instead of the explicit tags of FIGS. 4 and FIGS. 5.
[0066] As another example, the processor (220) can use the context of a previous or subsequent conversation based on the time the video was transmitted as metadata for the video in the case of a video among the videos owned by the user that was exchanged via messenger for clustering. In other words, keywords, summary sentences, and features generated as a result of analyzing the conversation segment including the time the video was transmitted can be added as information about the video.
[0067] FIGS. 6 to 8 illustrate an example of a conversation content analysis process in an embodiment of the present invention.
[0068] To analyze the content of a conversation, you can first identify the beginning and end of the conversation to select the section to be analyzed.
[0069] For example, the processor (220) can detect greetings indicating the start of a conversation or recognize a conversation segment based on a certain number of recent conversation messages. As another example, the processor (220) can divide the conversation segment into words that start after a specific period in time and words that end thereafter. Depending on the embodiment, a hybrid application based on the conversation content may also be applied. For example, if a conversation starts and then resumes after a 1-hour pause, and the conversation has the same topic as the previous conversation, it is also possible to combine the topics and analyze them at once.
[0070] If there are multiple conversation threads simultaneously, whether in a group chat or a one-on-one chat, the conversation context is analyzed to detect multiple topics, and different video content can be generated based on the conversation content.
[0071] The processor (220) can analyze the conversation content within the conversation section. The result of the conversation content analysis may include at least one topic and may have the form of keywords, summary sentences, features (embedding vectors), etc.
[0072] For example, the processor (220) can extract key keywords that are the subject of conversation from the conversation content using techniques such as graph ranking. In addition to graph ranking techniques, all natural language processing-based techniques for extracting key keywords from a document can be applied.
[0073] As another example, the processor (220) can summarize the conversation content using deep learning such as a large language model (LM). The content of the conversation segment to be analyzed can be summarized through a generative pre-trained transformer (chatGPT) to generate a summary sentence.
[0074] As another example, the processor (220) can generate an embedding for the conversation content. An embedding vector can be generated by converting the content of the conversation segment to be analyzed into text features through a text encoder. For the embedding vector, the entire content of the conversation segment to be analyzed can be converted into text features in bulk, and depending on the embodiment, it is also possible to summarize the conversation content and then convert the summary sentences into text features.
[0075] The aforementioned analysis of conversation content can be implemented in a form that is continuously performed whenever a conversation takes place, or by a method that is performed on the previous day's data during a rest period such as dawn.
[0076] Referring to FIGS. 6 to 8, the processor (220) can generate at least one of keywords, summary sentences, and features (embedding vectors) that can explain the topic of conversation well through continuous analysis of conversation content.
[0077] As illustrated in FIGS. 6 to 8, by analyzing conversation content over time, keywords, summary sentences, and features that can effectively explain the topic of conversation can be obtained, and based on this, related videos can be searched in the cluster DB to construct video content. Since this method generates necessary video content immediately, it is possible to recommend video content to be shared in real time.
[0078] Meanwhile, to reduce the computational burden of analyzing conversation content, the previous day's conversation content can be analyzed at a specific time (e.g., 3 AM) during a idle period when device usage is low, and keywords, summary sentences, and features that can effectively explain the topic of conversation for each conversation segment can be obtained.
[0079] FIG. 9 illustrates an example of a process for generating conversation-associated video content in an embodiment of the present invention.
[0080] Referring to FIG. 9, in step (S901), the processor (220) can search for a video having a similarity greater than or equal to a threshold value with a conversation topic, which is the result of analyzing conversation content, in the cluster DB for the video owned by the user. At this time, at least one of the keywords, summary sentences, and features obtained through the analysis of conversation content can be utilized for video search. At least one of the tags (keywords) and features obtained through video analysis can also be utilized for the video in the cluster DB.
[0081] The processor (220) can calculate the similarity between information obtained through conversation content analysis and information obtained through video analysis. For example, the processor (220) can calculate the text similarity between a keyword representing a conversation topic and a tag (keyword) obtained through video analysis. As another example, the processor (220) can calculate the similarity by converting a summary sentence representing a conversation topic into a text embedding and then comparing it with a feature (image embedding vector) obtained through video analysis. As yet another example, the processor (220) can calculate the similarity by converting a keyword representing a conversation topic into a text embedding and then comparing it with a feature (image embedding vector) obtained through video analysis.
[0082] The processor (220) selects a video that has a similarity to the result of analyzing the conversation content that is greater than or equal to a threshold value. The threshold value can be determined by using a predetermined value or by additionally training a machine learning model that selects the best video associated with the conversation content.
[0083] In step (S902), the processor (220) can generate video content for the conversation content using the video retrieved in step (S901), that is, the video whose similarity to the conversation content analysis result is greater than or equal to a threshold value. When using a single trigger, the generation of video content can begin when the condition is satisfied that at least one video whose similarity to the conversation content analysis result is greater than or equal to a threshold value is retrieved.
[0084] It is also possible to start creating video content using two or more triggers. By using situations such as when predefined keywords like 'photo' or 'share' are entered into a conversation message or when an album access menu is executed as additional triggers, video content creation can be started when an additional trigger situation is detected, along with a condition in which at least one video is found that has a similarity to the conversation content analysis result greater than or equal to a threshold value.
[0085] In step (S903), the processor (220) can recommend video content generated in step (S902).
[0086] For example, the processor (220) may periodically recommend video content generated in association with conversation content at a predetermined time or at a time set by the user. For example, as illustrated in FIG. 10, the processor (220) may provide a notification (1010) regarding conversation-related video content through the device screen (1000) at 9 a.m. every morning. At this time, the notification (1010) may include a link to play the conversation-related video content, and a photo application or messenger may be launched through the link so that the conversation-related video content can be played through the platform.
[0087] As another example, the processor (220) can recommend video content generated in association with conversation content in a form that can be shared via messenger. For example, as illustrated in FIG. 11, a notification (1110) regarding the sharing of conversation-related video content can be provided through the device screen (1000) at 9 a.m. every morning. At this time, the notification (1110) may include a link that allows the conversation-related video content to be shared on a predetermined platform, and through the link, a video-centric social content service function (e.g., LINE VOOM), which is one of the functions provided through the messenger platform, can be executed so that the conversation-related video content can be shared.
[0088] As another example, the processor (220) can recommend video content generated from the conversation content of the messenger chat room through the messenger chat room. If a conversation about a specific event is in progress, the processor can recommend content related to the event so that it can be sent. For example, as illustrated in FIG. 12, when a specific keyword (e.g., photo) is entered into the conversation message input window (1201) through the messenger chat room (1200), the processor can recommend conversation-related video content (1220) along with a sticker list (1210) that matches the keyword.
[0089] In this embodiment, video content generation is not triggered based on rules, but can be triggered by considering both the context of the conversation and the context of the user's video. By generating video content composed of photos of specific topics mentioned in the conversation rather than simply a video created with a theme such as 'one year ago today,' more timely and appropriate video content can be generated that matches the user's intent.
[0090] As such, according to the embodiments of the present invention, by triggering the generation of video content by considering both the context of the conversation in which the user is participating and the context of the photos held by the user, more meaningful video content can be generated, and the usability of photos that would otherwise only consume resources without any special use unless consumed by the user can be enhanced.
[0091] The device described above may be implemented as a hardware component, a software component, and / or a combination of a hardware component and a software component. For example, the device and components described in the embodiments may be implemented using one or more general-purpose or special-purpose computers, such as a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing unit may execute an operating system (OS) and one or more software applications executed on said operating system. Additionally, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing unit may be described as being used as a single unit, but those skilled in the art will understand that the processing unit may include multiple processing elements and / or multiple types of processing elements. For example, the processing unit may include multiple processors or one processor and one controller. In addition, other processing configurations, such as parallel processors, are also possible.
[0092] Software may include computer programs, code, instructions, or a combination of one or more of these, and may configure a processing unit to operate as desired or instruct the processing unit independently or collectively. Software and / or data may be embodied in any type of machine, component, physical device, computer storage medium, or device so as to be interpreted by the processing unit or to provide instructions or data to the processing unit. Software may be distributed over networked computer systems and may be stored or executed in a distributed manner. Software and data may be stored on one or more computer-readable recording media.
[0093] The method according to the embodiment may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. In this case, the medium may continuously store a program executable by a computer, or temporarily store it for execution or download. Additionally, the medium may be various recording or storage means in the form of a single or several hardware combined, and may not be limited to a medium directly connected to a computer system but may exist distributed over a network. Examples of media may include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and media configured to store program instructions, including ROM, RAM, and flash memory. Additionally, other examples of media may include recording or storage media managed by app stores that distribute applications or sites and servers that supply or distribute various other software.
[0094] Although the embodiments have been described above with reference to limited examples and drawings, those skilled in the art can make various modifications and variations from the description above. For example, suitable results can be achieved even if the described techniques are performed in a different order than described, and / or the components of the described system, structure, device, circuit, etc. are combined or assembled in a form different from described, or replaced or substituted by other components or equivalents.
[0095] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims set forth below.
Claims
1. A method for triggering image content generation of a computer device comprising at least one processor, A step of analyzing the conversation content of a messenger by the above-mentioned at least one processor; and A step of triggering the creation of video content using the searched video when at least one video among the user-held videos is searched by the at least one processor having a similarity value greater than or equal to a threshold value with respect to the analysis result of the conversation content. A method for triggering video content generation including 2. In Paragraph 1, The above-mentioned analysis step is, A step of determining the conversation segment to be analyzed; and Step of analyzing the conversation included in the above conversation section A method for triggering video content generation including 3. In Paragraph 1, The above-mentioned analysis step is, Step of extracting key keywords that serve as the topic of conversation from the above conversation content A method for triggering video content generation including 4. In Paragraph 1, The above-mentioned analysis step is, A step of generating a summary sentence by summarizing the above conversation content through a deep learning model. A method for triggering video content generation including 5. In Paragraph 1, The above-mentioned analysis step is, Step of converting the above conversation content into text embeddings through a text encoder A method for triggering video content generation including 6. In Paragraph 1, The above-mentioned analysis step is, A step of generating a summary sentence by summarizing the above conversation content through a deep learning model; and Step of converting the above summary sentence into text embeddings through a text encoder A method for triggering video content generation including 7. In Paragraph 1, The above video content generation triggering method is, A step of analyzing the user-held video and clustering it by topic by the at least one processor mentioned above A video content generation triggering method that further includes 8. In Paragraph 7, The above clustering step is, Generating a tag for each of the above user-owned videos and clustering the above user-owned videos by tag A video content generation triggering method characterized by 9. In Paragraph 7, The above clustering step is, Extracting features for each of the above user-owned images and clustering the above user-owned images based on the similarity between features A video content generation triggering method characterized by 10. In Paragraph 1, The above-mentioned triggering step is, A step of calculating the similarity between the information obtained through the analysis of the above conversation content and the information obtained through the analysis of the above user-held video; and A step of selecting an image with a similarity greater than or equal to the threshold value as an image associated with the conversation content. A method for triggering video content generation including 11. In Paragraph 1, The above-mentioned triggering step is, Triggering the creation of the video content when a video with a similarity greater than or equal to the above threshold value is found, and a predefined keyword or situation is detected. A video content generation triggering method characterized by 12. A computer program stored on a computer-readable recording medium to execute the video content generation triggering method of any one of claims 1 to 11 on the computer device.
13. At least one processor implemented to execute readable instructions on a computer device Includes, The above-mentioned at least one processor is, The process of analyzing the content of a messenger conversation; and A process of triggering the creation of video content using a found video when at least one video is found among the videos owned by the user that has a similarity value greater than or equal to a threshold value with respect to the analysis result of the above conversation content. A computer device that processes.
14. In Paragraph 13, The above-mentioned at least one processor is, Determine the conversation segments to be analyzed, and Analyzing the conversation included in the above conversation segment A computer device characterized by 15. In Paragraph 13, The above-mentioned at least one processor is, Obtaining at least one of a main keyword, a summary sentence, and a feature through the analysis of the above conversation content A computer device characterized by 16. In Paragraph 13, The above-mentioned at least one processor is, Analyzing the aforementioned user-owned videos and clustering them by topic A computer device characterized by 17. In Paragraph 16, The above-mentioned at least one processor is, Generating a tag for each of the above user-owned videos and clustering the above user-owned videos by tag A computer device characterized by 18. In Paragraph 16, The above-mentioned at least one processor is, Extracting features for each of the above user-owned images and clustering the above user-owned images through similarity between features A computer device characterized by 19. In Paragraph 13, The above-mentioned at least one processor is, Calculate the similarity between the information obtained through the analysis of the above conversation content and the information obtained through the analysis of the above user-owned video, and Selecting an image with a similarity greater than or equal to the threshold value as an image associated with the conversation content. A computer device characterized by 20. In Paragraph 13, The above-mentioned at least one processor is, Triggering the creation of the video content when a video with a similarity greater than or equal to the above threshold value is found, and a predefined keyword or situation is detected. A computer device characterized by