Program, method for generating digest videos, and information processing device.
The program and device automate the generation of digest videos using AI for video analysis, addressing the inefficiencies of human-dependent methods by predicting video contribution levels and reducing operational costs.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-18
- Publication Date
- 2026-03-31
AI Technical Summary
Generating digest videos from long videos requires human intervention based on past knowledge, making it costly and inefficient.
A program and information processing device that utilizes AI technologies for automatic video analysis, generating metadata, and predicting video contribution levels to support the generation of digest videos with reduced operational costs.
Enables efficient and cost-effective generation of digest videos by automating the process, improving processing efficiency and accuracy through machine learning and reinforcement learning techniques.
Smart Images

Figure 2026055236000001_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to a program, a method for generating a digest video, and an information processing apparatus.
Background Art
[0002] There is a technology that utilizes AI (Artificial Intelligence) technologies such as image recognition technology and speech recognition technology to automatically estimate various information contained in video data. For example, the following types of processing are possible.
[0003] Using person recognition technology, estimate the people appearing in the video and the times when they appear. Using object detection technology, estimate the objects that have moved in the video and the times when they have moved. Using OCR (Optical Character Recognition) technology, estimate the character strings in the video. Using speech recognition technology, convert the speech in the video into text.
[0004] Using proper noun extraction technology, extract proper nouns and keywords from character strings. Using speaker identification technology, estimate the speaker from the speech. Using emotion estimation technology, estimate the emotion from the speech. Using excitement estimation technology, estimate the range (time zone of the video) in which the speech is exciting.
[0005] Hereinafter, this video information obtained by automatic estimation is referred to as metadata. As an example of the use of this metadata, there is a video editing support apparatus that uses metadata. As application examples of the video editing support apparatus, for example, the following (1) and (2) can be considered.
[0006] (1) Extract from a long video only the parts that explain a specific product, generate a digest video of a specified length, and use the digest video for sales promotion. (2) Extract segments from long videos featuring popular individuals, create digest videos, and post them to video sharing sites. [Prior art documents] [Patent Documents]
[0007] [Patent Document 1] Japanese Patent Publication No. 2023-28548 [Overview of the Initiative] [Problems that the invention aims to solve]
[0008] However, generating digest videos from long videos like those described in (1) and (2) above was costly because it required human intervention based on past knowledge. Here, past knowledge refers to, for example, (3) and (4) below.
[0009] (3) If a retail store that sells products displays product description videos on monitors installed in the store, knowledge about how sales have changed when different types of videos have been displayed in the past. (4) If a digest video is posted to a video sharing site, knowledge regarding the number of viewers per unit of time, etc.
[0010] Therefore, it would be beneficial to provide support for generating digest videos at a low operational cost.
[0011] Therefore, the present invention has been made in view of the above circumstances, and aims to provide a program that can support the generation of digest videos related to products, a digest video generation method, and an information processing device. [Means for solving the problem]
[0012] The program of the embodiment causes the computer to function as follows: a first generation unit that performs analysis using image recognition technology and speech recognition technology on data of a first video for learning stored in a memory unit in association with a predetermined product, and generates first metadata including the analysis results; a contribution level identification unit that identifies a first video contribution level, which is the degree to which the first video contributed to the sales of the product, from the sales information of the product; a second generation unit that performs analysis using image recognition technology and speech recognition technology on data of a second video to be used for digest video generation, stored in the memory unit in association with the product, and generates second metadata including the analysis results; and a priority determination unit that predicts a second video contribution level, which is the degree to which the second video contributes to the sales of the product for each predetermined time period, based on the second metadata and a correlation model between the first metadata and the first video contribution level, and determines the priority for each time period from the second video contribution level. [Brief explanation of the drawing]
[0013] [Figure 1] A block diagram showing the functions of the information processing device according to the first embodiment. [Figure 2] A flowchart illustrating the learning process of an information processing device. [Figure 3] A flowchart illustrating the process of generating a digest video using an information processing device. [Figure 4A] An example of table data that maps videos to products. [Figure 4B] A diagram illustrating an example of information related to videos. [Figure 4C] A diagram showing an example of product information. [Figure 5] A diagram showing an example of metadata generated during training. [Figure 6] A diagram illustrating the processing logic of the learning unit. [Figure 7] A flowchart illustrating the process for determining priority. [Figure 8] A diagram showing an example of a UI for editing digest videos according to the second embodiment. [Figure 9]Explanatory diagram for generating video description text
Embodiments for Implementing the Invention
[0014] Hereinafter, with reference to the accompanying drawings, embodiments (First Embodiment, Second Embodiment) of the program, digest video generation method, and information processing apparatus of the present invention will be described. In the following, the term "assist" includes not only the case where a part of the work done by a person is taken over by a computer, but also the case where all of the work done by a person is taken over by a computer.
[0015] (First Embodiment) The information processing apparatus 1 (digest video generation support apparatus) of the first embodiment is used, for example, when generating a digest video of about 30 seconds for product explanation to be displayed on a monitor in a retail store selling products from a long video material of one hour or more.
[0016] The information processing apparatus 1 is a computer apparatus and includes a storage unit 2, an input unit 3, a display unit 4, a communication unit 5, and a processing unit 6.
[0017] The storage unit 2 is a storage device such as an HDD (Hard Disk Drive) or an SSD (Solid State Drive) and stores various types of information. The storage unit 2 stores, for example, video data 21, product sales information 22, metadata 23, a correlation model 24, etc.
[0018] In the following, among the videos, the video for learning is referred to as the "first video", and the video for which the digest video is to be generated is referred to as the "second video". Also, among the metadata, the metadata generated during learning is referred to as the "first metadata", and the metadata generated during digest video generation support is referred to as the "second metadata". Also, among the video contribution degrees, the video contribution degree calculated during learning is referred to as the "first video contribution degree", and the video contribution degree predicted during digest video generation support is referred to as the "second video contribution degree". Also, hereinafter, for the sake of simplicity of explanation, the descriptions of "first" and "second" may be omitted for the above terms.
[0019] The video data 21 includes data for a first video used for learning, which is stored in association with a predetermined product, and data for a second video, which is stored in association with a product and is the target of digest video generation. Note that the data for the first video and the data for the second video may each contain multiple products.
[0020] Here, Figure 4A shows an example of table data (video data information) that associates videos with products. As shown in Figure 4A, the association between videos and products is achieved by associating video IDs (Identifiers) with product IDs in the table data.
[0021] Figure 4B shows an example of information related to a video. As shown in Figure 4B, the registration date and video title are associated with the video ID as video attribute information. However, this is not limited to these; other information such as the creator and update date may also be associated. The information processing device 1 can search for videos based on this video attribute information.
[0022] Figure 4C shows an example of product information. As shown in Figure 4C, product information (product attribute information) includes the registration date, product title, and product category associated with the product ID. However, it is not limited to these, and other information may also be associated. The information processing device 1 can search for products based on this product attribute information.
[0023] Let's return to Figure 1 and continue the explanation. Product sales information 22 is information about past sales of a product. Alternatively, the product sales information 22 may be stored in an external device (not shown), and the information processing device 1 may retrieve and use the product sales information 22 from that external device.
[0024] Metadata 23 includes the first metadata and the second metadata (details below).
[0025] The correlation model 24 is created by the learning unit 64 (details will be described later).
[0026] The input unit 3 is an input device that receives user input to the information processing device 1, and is, for example, a keyboard, mouse, touch panel, etc.
[0027] The display unit 4 is a means for displaying various types of information, and is, for example, a display device such as an LCD (Liquid Crystal Display).
[0028] The communication unit 5 is a communication interface for communicating with external devices.
[0029] The processing unit 6 includes, for example, a CPU (Central Processing Unit), ROM (Read Only Memory), and RAM (Random Access Memory). The CPU comprehensively controls the operation of the information processing unit 1. ROM is a storage medium for storing various programs and data. RAM is a storage medium for temporarily storing various programs and for rewriting various data. The CPU then uses RAM as a work area to execute programs stored in ROM, storage unit 2, etc.
[0030] The processing unit 6 comprises, functionally, an acquisition unit 61, a generation unit 62, a contribution level identification unit 63, a learning unit 64, a priority determination unit 65, a digest video generation unit 66, and a control unit 67.
[0031] The acquisition unit 61 acquires various types of information from external devices, storage unit 2, input unit 3, etc.
[0032] The generation unit 62 performs various generation processes. For example, the generation unit 62 (first generation unit) analyzes the data of the first video for learning, which is stored in the storage unit 2 in association with a predetermined product, using image recognition technology and speech recognition technology, and generates first metadata including the analysis results. The generated first metadata is stored in the storage unit 2 as metadata 23.
[0033] Each first metadata entry includes an attribute indicating the analysis category, an attribute value indicating the estimation result, and a start time and end time indicating the estimation time. It may also include the confidence level of the analysis. Figure 5 shows an example of metadata generated during training. From left to right, each item is: start time indicating the estimation time, end time, attribute, attribute value, and analysis confidence level.
[0034] The generation unit 62 generates the first metadata by, for example, the following process: Using person recognition technology, we estimate the characters in the video and the time they appear. Using object detection technology, we estimate the objects that appear in the video and the time at which they appeared. OCR technology is used to estimate text within a video. Speech recognition technology is used to convert the audio in a video into text.
[0035] This technique uses proper noun extraction technology to extract proper nouns and keywords from a string of text. Speaker estimation technology is used to estimate the speaker from the audio. Note that the speaker may or may not be visible in the video. Emotion estimation technology is used to estimate emotions from speech. Using excitement estimation technology, the range of excitement (time period in the video) is estimated from the audio.
[0036] Returning to Figure 1, let's continue the explanation. Alternatively, the generation unit 62 may use scene estimation technology to divide the video into scenes and generate metadata summarizing each analysis result for each scene based on the time information of each analysis result.
[0037] In the following, explanations of the second metadata will be omitted as appropriate for matters similar to those of the first metadata. For example, the generation unit 62 (second generation unit) performs analysis using image recognition technology and speech recognition technology on the data of the second video targeted for digest video generation, which is stored in association with the product in the storage unit 2, and generates the second metadata including the analysis results.
[0038] In that case, the generation unit 62 (second generation unit) may, when generating the second metadata, generate only metadata whose correlation with the first video contribution is greater than or equal to a predetermined correlation threshold, based on a correlation model. This would allow for a reduction in processing time without significantly reducing the accuracy of the processing results, thus improving processing efficiency.
[0039] The contribution determination unit 63 calculates the first video contribution, which is the degree to which the first video contributed to the sales of the product, from the product sales information.
[0040] For example, the contribution identification unit 63 obtains "sales information when a video introducing the product is shown in the sales area" (sales information A) and "sales information when a video introducing the product is not shown in the sales area" (sales information B), and calculates the video contribution as follows. Video contribution = Sales information A / Sales information B
[0041] The contribution level identification unit 63 may, for example, obtain a pre-calculated first video contribution level from an external device (not shown).
[0042] The learning unit 64 learns the correlation between the first video contribution calculated by the contribution identification unit 63 and the first metadata generated by the generation unit 62 (first generation unit) and creates a correlation model. The created correlation model is stored in the storage unit 2 as the correlation model 24.
[0043] Specifically, the learning unit 64 learns the correlation between the attribute values of the first metadata as input values and the first video contribution as output values. Conventional machine learning techniques such as deep learning are used for learning. For example, the model is as follows. Here, Figure 6 is a diagram illustrating the processing image of the learning unit 64.
[0044] y(i) = W * M(i,n) M(i,n)=concat(vector(i,n,k))
[0045] Hereinafter, we will proceed as follows: y(i): The first video contribution of video i. W: Learning weights (correlation between the first video contribution and the first metadata) (Figure 6(b)) M(i,n): Vector of the first metadata for frame n of video i vector(i,n,k): A vector of attribute k of frame n in video i.
[0046] The vector M(i,n) is generated by concatenating the vectors of each attribute for each unit of time (frame) (Figure 6(a)). The vector vector(i,n,k) of each attribute uses the average or maximum value per unit of time.
[0047] Alternatively, the video contribution y(i) may be calculated by dividing it by the total number of frames. Alternatively, the video contribution y(i) of the previous frame may be added to M(i,n). Furthermore, attributes of video data information (for example, the product category) can be added to M(i,n).
[0048] Furthermore, the learning unit 64 may be reinforcement learning, which is a type of machine learning that uses rewards, and may create a correlation model by performing reinforcement learning that reflects actual sales information in the rewards. Specifically, it is as follows:
[0049] The learning unit 64 is configured to learn correlations using reinforcement learning technology that uses actual sales information as a reward for the videos generated by the digest video generation unit 66. As a result, the learning unit 64 can learn the contribution of the first video more accurately based on actual sales information, and improve the accuracy of predicting the contribution of the second video in the future.
[0050] Reinforcement learning is a method in which the learning unit 64 (agent) learns actions that maximize rewards while interacting with the environment. For example, the process of using actual sales information as a reward is as follows:
[0051] 1. Initial Setup The learning unit 64 acquires sales information for the video generated by the digest video generation unit 66.
[0052] 2. Interaction with the environment The learning unit 64 obtains the first metadata used by the digest video generation unit 66 when generating the video, and the first video contribution score.
[0053] 3. Choosing an action The learning unit 64 selects the next video data to analyze based on either the first video contribution, sales information, or both. For example, it selects video data with a high first video contribution.
[0054] 4. Receiving rewards Based on actual sales data, the learning unit 64 calculates the rewards obtained as a result of its actions. For example, if a particular video contributes significantly to sales, it assigns a high reward to the metadata associated with that video.
[0055] 5. Policy Updates The learning unit 64 updates the correlation based on the rewards obtained.
[0056] 6. Repeat The learning unit 64 repeatedly performs the above process and gradually learns the optimal course of action.
[0057] In this way, the learning unit 64 can reflect actual sales information as a reward and improve the prediction accuracy of the second video contribution through reinforcement learning.
[0058] The priority determination unit 65 predicts the second video contribution, which is the degree to which the second video contributes to product sales, for each predetermined time period of the second video, based on the second metadata and the correlation model 24 between the first metadata and the first video contribution, and determines the priority for each time period from the second video contribution. The predetermined time period may be, for example, each frame, each time period of about one second, or each scene.
[0059] Here, Figure 7 is a flowchart showing the process when determining priority. In step S31, the priority determination unit 65 obtains the correlation model 24 from the storage unit 2.
[0060] Next, in step S32, the priority determination unit 65 obtains product information for which a digest video is to be generated.
[0061] Next, in step S33, the priority determination unit 65 determines the range related to product information (related range) from the second metadata generated by the generation unit 62. For example, it searches for second metadata whose attribute is product estimation and whose attribute value includes the ID or title of the target product. Alternatively, it searches for metadata whose attribute is speech recognition result and whose attribute value includes the title of the product. Alternatively, it searches for metadata whose attribute is caption recognition result and whose attribute value includes the title of the target product.
[0062] Then, the range (related range) is determined from the start and end times of these second metadata. For example, the range start time is set by subtracting a certain amount of time from the minimum start time of the second metadata, and the range end time is set by adding a certain amount of time to the maximum end time of the second metadata.
[0063] Next, in step S34, the priority determination unit 65 divides the second metadata of the relevant range into frame units.
[0064] Next, in step S35, the priority determination unit 65 vectorizes the second metadata on a frame-by-frame basis.
[0065] Next, in step S36, the priority determination unit 65 predicts a second video contribution on a frame-by-frame basis from the frame-by-frame vector and the correlation model 24.
[0066] Next, in step S37, the priority determination unit 65 increases the priority of frames with a large second video contribution. At the same time, the priority of frames adjacent to frames with a large second video contribution may also be increased.
[0067] Returning to Figure 1, the explanation continues. The digest video generation unit 66 extracts the range of the digest video from the data of the second video and generates the digest video based on the priority determined by the priority determination unit 65.
[0068] Specifically, for example, the digest video generation unit 66 limits the total time of the digest video, adds frames in order of priority, and extracts frames that fall below the specified total time. The digest video generation unit 66 then extracts frames from the beginning, concatenates the extracted frames in order, and generates a digest video.
[0069] The control unit 67 performs various controls.
[0070] Next, referring to Figures 2 and 3, we will explain the flow for learning the correlation between the first metadata and the first video contribution using one or more videos related to a product (Figure 2), and the flow for generating a digest video for any given video after learning (Figure 3).
[0071] Figure 2 is a flowchart showing the processing of the information processing device 1 during learning. Here, it is assumed that in the storage unit 2, the data of the first video in the video data 21 is stored in association with a predetermined product.
[0072] In step S11, the generation unit 62 (first generation unit) performs analysis on the data of the first video using image recognition technology and speech recognition technology, and generates first metadata including the analysis results.
[0073] Next, in step S12, the contribution identification unit 63 calculates a first video contribution from the product sales information 22, which is the degree to which the first video contributed to the sales of the product.
[0074] Next, in step S13, the learning unit 64 learns the correlation between the first video contribution calculated in step S12 and the first metadata generated in step S11, and creates a correlation model 24.
[0075] Next, Figure 3 is a flowchart showing the processing when the information processing device 1 generates a digest video. Here, it is assumed that in the storage unit 2, the data for the second video in the video data 21 is stored in association with a predetermined product.
[0076] In step S21, the generation unit 62 (second generation unit) performs analysis on the data of the second video to be used for digest video generation using image recognition technology and speech recognition technology, and generates second metadata including the analysis results.
[0077] Next, in step S22, the priority determination unit 65 predicts the second video contribution, which is the degree to which the second video contributes to product sales, for each predetermined time period of the second video, based on the second metadata and the correlation model 24 between the first metadata and the first video contribution.
[0078] Next, in step S23, the priority determination unit 65 determines the priority for each time period based on the second video contribution predicted in step S22.
[0079] Next, in step S24, the digest video generation unit 66 extracts the range of the digest video from the data of the second video based on the priority determined in step S23 and generates the digest video.
[0080] Thus, according to the information processing device 1 of the first embodiment, based on the second metadata and the correlation model 24 between the first metadata and the first video contribution, the second video contribution, which is the degree to which the second video contributes to product sales, is predicted for each predetermined time period of the second video, and the priority for each time period is determined from the second video contribution. This enables support for generating digest videos related to products.
[0081] Furthermore, the learning unit 64 can create a highly accurate correlation model 24 by learning the correlation between the first metadata and the first video contribution.
[0082] Furthermore, the digest video generation unit 66 can automatically generate a digest video by extracting the range of the digest video from the data of the second video based on priority.
[0083] Furthermore, if the generation unit 62 (second generation unit) generates the second metadata, and based on the correlation model 24, it generates only metadata whose correlation with the first video contribution is greater than or equal to a predetermined correlation threshold, the processing can be made more efficient.
[0084] Furthermore, if the learning unit 64 performs reinforcement learning, which is a type of machine learning using rewards, and if it reflects actual sales information in the rewards, a more accurate correlation model 24 can be created.
[0085] (Second Embodiment) Next, a second embodiment will be described. Matters similar to those in the first embodiment will be omitted from the explanation as appropriate. The second embodiment differs from the first embodiment in that it has a user interface (UI) for editing videos. Figure 8 shows an example of the UI for editing digest videos in the second embodiment.
[0086] The digest video generation unit 66 extracts the range of the digest video from the data of the second video and performs the following processing when generating the digest video.
[0087] The digest video generation unit 66 performs control to play and stop the video in the video playback area (area R1 in Figure 8) that displays the data of the second video, in response to user operations.
[0088] The digest video generation unit 66 executes control to select a section specified by the user (in the example in Figure 8, the section for product A) in the section selection area (area R2 in Figure 8) that displays the section corresponding to the product estimated based on the second metadata for the second video. The estimated section displayed in the section selection area (area R2 in Figure 8) can be, for example, a certain section before and after the time when the product was detected using the object detection result or speech recognition result of the second metadata.
[0089] The digest video generation unit 66 performs control to select one or more recommended sections specified by the user in the digest area (area R3 in Figure 8), where the priority determined by the priority determination unit 65 within a section (the section for product A in the example of Figure 8) is equal to or greater than a predetermined priority threshold, and displays these sections as recommended sections (P1, P2, P3).
[0090] The digest video generation unit 66 executes control to display the second video contribution of the recommended section (e.g., P1, P2, P3) selected in the digest area (area R5 in Figure 8) in the digest information presentation area (area R3 in Figure 8).
[0091] The digest video generation unit 66 executes control to display second metadata (objects, people, emotions, captions, excitement, etc.) in the metadata display area (area R4 in Figure 8) for the section (the section of product A in area R2 in Figure 8).
[0092] The digest video generation unit 66 extracts the selected recommended portion (e.g., P1, P2, P3) from the data of the second video within the digest region (region R3 in Figure 8) and executes control to generate a digest video.
[0093] Furthermore, the priority determination unit 65 may store each of the above-mentioned user operations as an editing history in the storage unit 2, and then, when determining the priority for the third video (the video to be used to generate a new digest video), it may also use the editing history to determine the priority. Specifically, for example, it may be done as follows.
[0094] The section selected in the digest area of the user interface (area R3 in Figure 8) is saved as the editing history. Then, the second metadata for the section saved in the editing history is retrieved. Next, when determining the priority of the third video, if the third metadata calculated for the third video matches or is similar to its second metadata, the calculated contribution of the third video is increased by a certain amount.
[0095] Next, Figure 9 is an explanatory diagram for generating a video description. The digest video generation unit 66 may also perform control to display the acquired second video description by inputting second metadata corresponding to the recommended portion selected in the digest region (region R3 in Figure 8) into the generation AI.
[0096] In the example in Figure 9, a general generative AI technique may be used to generate a prompt for the generative AI based on the second metadata of the recommended section selected in the digest area of the user interface (area R3 in Figure 8) (Figure 9(a)), input the generated prompt into the generative AI to create a description of the digest video (Figure 9(b)), and then present the created description in the digest information presentation area (area R5 in Figure 8).
[0097] Thus, according to the second embodiment, the generation of a digest video can be supported in a way that allows the user to select the parts of the second video that they want to use as a digest video. Therefore, the digest video can be made to contain content that is more desirable to the user.
[0098] Furthermore, the information processing device 1 described above can also be realized, for example, by using a general-purpose computer device as the basic hardware. That is, each part 61 to 67 of the processing unit 6 can be realized by having a processor installed in the information processing device 1 execute a program. In this case, the information processing device 1 may be realized by pre-installing the above program on the computer device, or by storing the above program on a storage medium such as a CD-ROM, or by distributing the above program via a network and installing this program on the computer device as appropriate.
[0099] Furthermore, the storage unit 2 can be implemented using memory, a hard disk, or storage media such as a CD (Compact Disc)-R (Recordable), CD-RW, DVD (Digital Versatile Disk)-RAM, or DVD-R, which are built into or attached to the information processing device 1.
[0100] While several embodiments of the present invention have been described, these embodiments are presented as examples only and are not intended to limit the scope of the invention. These novel embodiments can be carried out in a variety of other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their variations are included in the scope and spirit of the invention, as well as in the claims of the invention and its equivalents.
[0101] For example, in the above-described embodiment, the information processing device 1 was described as being composed of a single computer device for the sake of simplicity, but it is not limited to this. The information processing device 1 may be implemented by, for example, multiple computer devices. Furthermore, at least a part of the functions of the information processing device 1 may be implemented by cloud computing. [Explanation of Symbols]
[0102] 1...Information processing unit, 2...Storage unit, 3...Input unit, 4...Display unit, 5...Communication unit, 6...Processing unit, 21...Video data, 22...Product sales information, 23...Metadata, 24...Correlation model, 61...Acquisition unit, 62...Generation unit, 63...Contribution level identification unit, 64...Learning unit, 65...Priority determination unit, 66...Digest video generation unit, 67...Control unit
Claims
1. Computers, A first generation unit performs analysis on the data of a first video for learning, which is stored in the memory unit in association with a predetermined product, using image recognition technology and speech recognition technology, and generates first metadata including the analysis results. A contribution level identification unit identifies a first video contribution level, which is the degree to which the first video contributed to the sales of the product, based on the sales information of the product, A second generation unit performs analysis using image recognition technology and speech recognition technology on the data of a second video to be used for digest video generation, which is stored in the storage unit in association with the product, and generates second metadata including the analysis results. A program to function as a priority determination unit that predicts the degree to which the second video contributes to the sales of the product for each predetermined time period of the second video, based on the second metadata and the correlation model between the first metadata and the first video contribution, and determines the priority for each time period from the second video contribution.
2. The aforementioned computer, Furthermore, the program according to claim 1, which functions as a learning unit that learns the correlation between the first metadata and the first video contribution and creates the correlation model based on the first video contribution identified by the contribution identification unit and the first metadata generated by the first generation unit.
3. The aforementioned computer, Furthermore, the program according to claim 1, which functions as a digest video generation unit that extracts a range of the digest video from the data of the second video and generates the digest video based on the priority determined by the priority determination unit.
4. The program according to claim 1, wherein the second generation unit, when generating the second metadata, generates only metadata as the second metadata whose magnitude of correlation with the first video contribution is equal to or greater than a predetermined correlation threshold, based on the correlation model.
5. The program according to claim 2, wherein the learning unit is a reinforcement learning machine learning using rewards, and the program creates the correlation model by performing the reinforcement learning which reflects actual sales information in the rewards.
6. The aforementioned computer, Furthermore, a digest video generation unit extracts a range for a digest video from the data of the second video mentioned above and generates a digest video, In the video playback area that displays the data of the second video, control is provided to play and stop the video in response to user operation, With respect to the second video, in the section selection area that displays the section corresponding to the product estimated based on the second metadata, control is provided to select the section specified by the user operation, Within the aforementioned section, in a digest area where the priority determined by the priority determination unit is equal to or greater than a predetermined priority threshold is displayed as a recommended section, control is provided to select one or more of the recommended sections specified by the user operation. Control to display the second video contribution of the recommended portion selected in the digest area within the digest area, For the aforementioned section, control is provided to display the second metadata in the metadata display area, The program according to claim 1, which functions as a digest video generation unit that performs control to extract the range of the recommended portion selected in the digest region from the data of the second video and generate the digest video.
7. The program according to claim 6, wherein the priority determination unit stores the user operations as editing history in the storage unit, and thereafter, when determining the priority for the third video, the program also uses the editing history to determine the priority.
8. The aforementioned digest video generation unit, The program according to claim 6, which performs control to display the description of the second video obtained by inputting the second metadata corresponding to the recommended portion selected in the digest area into the generating AI.
9. The first generation step involves the first generation unit performing analysis on the first video data for learning, which is stored in the memory unit in association with a predetermined product, using image recognition technology and speech recognition technology, and generating first metadata including the analysis results. The contribution identification unit identifies a first video contribution, which is the degree to which the first video contributed to the sales of the product, from the sales information of the product. The second generation step involves the second generation unit performing analysis on the data of the second video to be used for digest video generation, which is stored in the storage unit in association with the product, using image recognition technology and speech recognition technology, and generating second metadata including the analysis results. Priority determination step: A priority determination unit predicts the second video contribution, which is the degree to which the second video contributes to the sales of the product, for each predetermined time period of the second video, based on the second metadata and the correlation model between the first metadata and the first video contribution, and determines the priority for each time period from the second video contribution. A digest video generation method comprising a digest video generation step in which a digest video generation unit extracts a range for a digest video from the data of the second video based on the priority determined by the priority determination step, and generates a digest video.
10. A first generation unit performs analysis on the data of a first video for learning, which is stored in the memory unit in association with a predetermined product, using image recognition technology and speech recognition technology, and generates first metadata including the analysis results. A contribution level identification unit identifies a first video contribution level, which is the degree to which the first video contributed to the sales of the product, based on the sales information of the product, A second generation unit performs analysis using image recognition technology and speech recognition technology on the data of a second video to be used for digest video generation, which is stored in the storage unit in association with the product, and generates second metadata including the analysis results. An information processing device comprising: a priority determination unit that predicts a second video contribution, which is the degree to which a second video contributes to the sales of the product, for each predetermined time period of the second video, based on the second metadata and a correlation model between the first metadata and the first video contribution, and determines the priority for each time period from the second video contribution.
Citation Information
Patent Citations
Information processing device, information processing method, and information processing program
JP2023028548A