Support device for editing multimedia content

JP7899429B1Active Publication Date: 2026-08-03ZEAL CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
ZEAL CO LTD
Filing Date
2025-09-30
Publication Date
2026-08-03

AI Technical Summary

Benefits of technology

【0016】 本発明は、閲覧者に喚起する感情の進行に沿った変化を踏まえた複合媒体コンテンツの編集を支援する技術的手段を提供できる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007899429000001_ABST
    Figure 0007899429000001_ABST
Patent Text Reader

Abstract

To support the editing of multimedia content that takes into account the progression of emotions evoked in the viewer. [Solution] The composite media content editing support device 1 of the present invention comprises: a content data acquisition unit 111 that acquires content data to be used for displaying content including multiple types of media; an analysis unit identification unit 112 that identifies analysis units in line with the progress of the content; an emotion information acquisition unit 113 that inputs data corresponding to the analysis unit based on content data to a converter for each analysis unit and outputs emotion information including the intensity of multiple types of emotions evoked by a part of the content related to the analysis unit; and an analysis information generation unit 114 that generates analysis information related to content editing based on the changes in the emotion index calculated based on this emotion information in line with the progress and criteria related to the emotion index, and inputs data showing these changes to a large-scale language model to generate text related to the analysis information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an apparatus for assisting in the editing of composite media content.

Background Art

[0002] As a technique for creating content that appeals more deeply to the emotions of viewers and gives a stronger impression, there is a technique of creating content based on changes along the progression of the emotions evoked by the content in the viewer. If the changes can be analyzed, such creation can be carried out more efficiently or effectively. Therefore, in the process of creating content, technical means for assisting the analysis of changes along the progression of the emotions evoked by the content in the viewer are required.

[0003] As means for assisting the analysis of the above changes, Patent Document 1 discloses an article analyzer including an article acquisition unit, an article segmentation unit, an emotion vector generation unit, an evaluation generation unit, a display command unit, and a first machine learning unit, and the evaluation generation unit can generate an evaluation corresponding to a correlation value regarding at least the correlation between the norm of the emotion vector and the moving average of the norm.

[0004] The technique of Patent Document 1 can provide an analysis result that can easily grasp the characteristics of an article.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0006] Incidentally, not only single-media content such as text, but also multimedia content that includes multiple different types of media as constituent elements is widely produced and enjoyed. Examples of multimedia content include "manga" which includes image and text media, "videos" which include video and audio media, and "multimedia posts" on social networking services (SNS) which include at least two of the following: image, video, audio, and text media.

[0007] In multimedia content, unlike single-media content, not only are the emotions evoked by each individual medium in the content present in the viewer, but the emotions evoked by combining these media can also change. Furthermore, in content where images, videos, and other non-textual media are the primary elements, analyzing processing units divided based on the primary elements, rather than using processing units based on text, allows for a more appropriate analysis of the emotions evoked by the content in the viewer.

[0008] While the technology described in Patent Document 1 provides analysis results that allow for easy understanding of the characteristics of text, there is room for further improvement in analyzing changes in line with the progression of emotions evoked by multimedia content in the viewer and providing editing support for said content based on that analysis. For these reasons, there is a need for technical means to support the editing of multimedia content that takes into account changes in line with the progression of emotions evoked in the viewer.

[0009] This invention was made to solve the problems of the prior art described above, and aims to provide a technical means to support the editing of multimedia content that takes into account changes in line with the progression of emotions evoked in the viewer. [Means for solving the problem]

[0010] As a result of diligent research to solve the above problems, the inventors have found that the above problems can be solved by using a large-scale language model in conjunction with a configuration that generates analytical information related to editing based on changes in sentiment indicators in line with the progression and their criteria, and by other technical means. The inventors have now completed the present invention. Specifically, the present invention provides the following:

[0011] An invention according to one aspect of the present invention includes: a content data acquisition unit that acquires content data to be used for displaying content including multiple types of media; an analysis unit identification unit that identifies analysis units in line with the progression of the content based on the content data; an emotion information acquisition unit that inputs data corresponding to the analysis unit based on the content data to a converter for each analysis unit and outputs emotion information including the intensity of multiple types of emotions evoked by a part of the content relating to the analysis unit; and an analysis information relating to the editing of the content based on the changes in the emotion index calculated based on the emotion information in line with the progression and criteria relating to the emotion index. Analytical values The analysis information generation unit inputs the data into a large-scale language model to generate text related to the analysis information, and The aforementioned analytical values ​​show changes in line with the aforementioned progression. We provide a support device for editing multimedia content.

[0012] The emotion information acquisition unit of the present invention inputs data corresponding to analysis units in line with the progression of the content, which have been identified based on the content data, into a converter, and outputs emotion information including the intensity of multiple types of emotions evoked by a part of the content relating to the analysis unit. The analysis information generation unit of the present invention then generates analysis information related to the editing of the content based on the changes in the emotion index calculated based on the emotion information in line with the progression and criteria related to the emotion index.

[0013] At this time, the analysis information generation unit of the invention inputs data showing changes in line with the progression of the emotion index into a large-scale language model to generate text related to the analysis information. The large-scale language model typically generates text based on trained parameters that represent the contextual relationships of the input token sequence. Therefore, the invention can generate text based on the large-scale language model that includes not only the emotions evoked in the viewer by each medium constituting the content, but also descriptions of the emotions evoked in the viewer due to the combination of each medium.

[0014] As a result, the invention can provide a suggestion function for editing multimedia content based on multimodal sentiment analysis of said content. Therefore, the invention according to this aspect of the present invention can provide a technical means to support the editing of multimedia content that takes into account changes in line with the progression of emotions evoked in the viewer.

[0015] Furthermore, the present invention can take various forms, including an embodiment that enables the identification of analysis units in a wide range of videos by using multiple analysis unit identification criteria, an embodiment that addresses the problem that identification based on the magnitude of changes in a video is not always successful by using fallback processing of the analysis unit identification means, and others. Each of these embodiments of the present invention contributes to supporting the editing of multimedia content that takes into account changes in line with the progression of emotions evoked in the viewer, each with its own unique characteristics. [Effects of the Invention]

[0016] This invention provides technical means to support the editing of multimedia content that takes into account changes in line with the progression of emotions evoked in the viewer. [Brief explanation of the drawing]

[0017] [Figure 1] Figure 1 is a block diagram showing an example of the hardware and software configuration of system S. [Figure 2] Figure 2 shows an example of the analytical information database 131. [Figure 3]FIG. 3 is a main flowchart showing an example of a preferable flow of assistance processing executed by the assistance device 1 of the present embodiment. [Figure 4] FIG. 4 is a figure following the previous figure. [Figure 5] FIG. 5 is a figure following the previous figure. [Figure 6] FIG. 6 is an example of a display related to analysis information. [Figure 7] FIG. 7 is an example of a detailed display related to analysis information of a video. [Figure 8] FIG. 8 is an example of a detailed display related to analysis information of text.

MODE FOR CARRYING OUT THE INVENTION

[0018] First of all, although the following disclosure, charts, and / or claims, etc. are described as being alone or in combination with one or more other aspects, the subject matter of the instant disclosure is not intended to be limited in that way. That is, the instant disclosure, charts, and claims are intended to encompass the various aspects described herein, either alone or in one or more combinations with each other. For example, even if the instant disclosure describes and illustrates the first embodiment, the second embodiment, and the third embodiment in such a way that the first embodiment is particularly described and illustrated in relation to the second embodiment, or the second embodiment is only described and illustrated in relation to the third embodiment, the instant disclosure and illustration are not limited in that way, and may include only the first embodiment, only the second embodiment, only the third embodiment, or one or more combinations of the first, second, and / or third embodiments, for example, the first embodiment and the second embodiment, the first embodiment and the third embodiment, the second embodiment and the third embodiment, or the first, second, and third embodiments.

[0019] The use of the phrase "or" in this document shall mean a "non-exclusive" determination, unless otherwise specified. For example, in the case of "Item x is A or B", it shall mean either of the following: (1) Item x is only one of A or B, (2) Item x is both A and B. In other words, the word "or" is not used to define an "exclusive" determination.

[0020] Also, when the phrases "including at least one" or "including at least one of the following" are used in this document and combined with a system or element, it means that the system or element includes one or more of the elements listed after the phrase. For example, when there are three types of elements from the first element to the third element, the phrases "including at least one" or "including at least one of the following" shall be interpreted as any of the following structural arrangements: a device including the first element, a device including the second element, a device including the third element, a device including the first and second elements, a device including the first and third elements, a device including the second and third elements, or a device including the first, second, and third elements.

[0021] The same interpretation is intended when the phrase "used in at least one of the following" is used in this document. Furthermore, "and / or" used in this document is used as a linguistic conjunction and is used to indicate that one or more of the described elements or conditions are included or occur. For example, a device including the first element, the second element, and / or the third element shall be interpreted as any of the following structural arrangements: a device including the first element, a device including the second element, a device including the third element, a device including the first and second elements, a device including the first and third elements, a device including the second and third elements, or a device including the first, second, and third elements.

[0022] Note that the fact that the use of the phrase "and / or" in this document means a "non-exclusive" determination is also stipulated in the "Format and Preparation Method of Standard Sheets JIS Z 8301" of the Japanese Industrial Standards (JIS).

[0023] Hereinafter, an example of an embodiment of the present invention will be described in detail with reference to the drawings.

[0024] <System S> Figure 1 is a block diagram showing an example of the hardware and software configuration of System S.

[0025] System S comprises a support device (support device 1) for editing composite media content. The following describes a server-client configuration in which support device 1 collaborates with terminals T configured to communicate with each other via a network N. Those skilled in the art will readily conceive of a standalone configuration of support device 1, including input and display means, based on the following description.

[0026] [Support Device 1] Support device 1 according to this embodiment is a computer comprising a control unit 11, a storage unit 13, and a communication unit 14. Support device 1 performs support processing for editing composite media content by executing the program of this embodiment.

[0027] [Control Unit 11] The control unit 11 includes a Central Processing Unit (CPU), Random Access Memory (RAM), Read Only Memory (ROM), and other hardware components.

[0028] The control unit 11 cooperates with at least one of the storage unit 13 and the communication unit 14 as needed. The control unit 11 then implements the content data acquisition unit 111, the analysis unit identification unit 112, the emotion information acquisition unit 113, the analysis information generation unit 114, and the visual representation data generation unit 115, which are software components of the program of this embodiment executed by the support device 1.

[0029] The support processing performed by the program of this embodiment will be described later with reference to Figures 3 to 5.

[0030] [Storage Unit 13] The storage unit 13 is a device on which data and / or files are stored, and has a storage unit that stores data non-temporarily using a hard disk, semiconductor memory, recording medium, memory card, or other storage material. Preferably, the storage unit 13 stores data for programs executed by the control unit 11, an analysis information database 131, content data, and other data.

[0031] Furthermore, in order to achieve generation using a large-scale language model without transmitting content data to an external device that is different from the support device 1 and the managing entity, the storage unit 13 may store data for executing the large-scale language model according to this embodiment.

[0032] (Analysis Information Database 131) The analysis information database 131 stores the analysis information generated by the support device 1. The analysis information stored in the analysis information database 131 includes content data (e.g., video, images, audio, text), non-analysis result data based on the content data (e.g., speech recognition results), various sentiment information and sentiment indicators corresponding to the analysis units identified based on the content data, and analysis information text generated based on these.

[0033] The analysis information database 131 stores, for example, information identifying the time or sequence corresponding to the analysis unit (hereinafter referred to as "progress information"), the emotion vector related to the analysis unit (e.g., intensity values ​​of Plutchik's eight basic emotions), the minimum and maximum values ​​based on each element of the emotion vector, the emotion value based on the emotion vector (e.g., normalized Euclidean norm), the moving average of the emotion value, the correlation coefficient between the emotion value and the moving average, and other indicators. Preferably, the analysis information database 131 further stores content characteristics based on the analysis unit, editing advice, summaries, and other analysis information text.

[0034] In this way, by storing analytical information including sentiment information, sentiment indicators, analytical information text, and other data in the analytical information database 131, the support device 1 can visually represent the changes in sentiment in line with the progression of the content in a graph format during subsequent display processing, and present comparisons with standards and advice to the user. Furthermore, the analytical information database 131 can also be configured to accumulate the results of multiple analyses corresponding to the edited version, making it possible to track the effects of editing improvements.

[0035] More detailed information about the analysis will be provided in the explanation of the support processing flow described later. In the following, the analysis information will be described as being stored in association with an identification ID (analysis information ID) used for storage, retrieval, and other data utilization.

[0036] Figure 2 shows an example of the analysis information database 131. This example shows the first analysis information relating to the entire video content "The Secret Story of Anpan's Birth," associated with analysis information ID "A0001," and the second analysis information relating to a single analysis unit of the same video content, associated with analysis information ID "A1002." For the sake of clarity, analysis information relating to other analysis units has been omitted from the diagram.

[0037] The first analysis information for this example includes the content's overall score (60 points) and its ratio to the benchmark (out of 100 points). This information also includes the final evaluation "Rank C" based on the overall score and text generated by a large-scale language model ("To approach the 'man in a hole' type, it's important to first set a high baseline for the protagonist at the beginning of the story, then deeply depict the emotional 'valley.' And the frustration..."). Furthermore, this information includes graph correlation: 0.33 and its satisfaction rate against the benchmark of 30%, average emotion: 51.2 and its satisfaction rate against the benchmark of 100%, and emotion change rate: 0.58 and its satisfaction rate against the benchmark of 50%.

[0038] Because the example includes the first analysis information described above, the support device 1 can comprehensively analyze the changes in the entire video content in line with the progression of emotions, and present the analysis results to the user as numerical values ​​(scores), ranks, and generated text. As a result, the user can intuitively grasp the overall completeness and direction of improvement of the content, and determine the priorities for editing.

[0039] The second analysis information for this example includes a representative frame of one analysis unit (scene) of the video content ("a frame depicting a man surprised by Western culture") and the text of the speech recognition results of the audio contained in the video content ("Around the time of the Meiji Restoration, those who liked Western culture were sometimes ridiculed as 'Westernized,' but those who couldn't keep up with the times were said to be 'Tenpo old men'..."). This information also includes editing advice generated by a large-scale language model ("Let's make the characters' expressions and reactions richer to emphasize their surprise and bewilderment at Western culture. Let's use backgrounds and sound effects to mask the uncertainty of the rumors and encourage the reader to empathize. 'Anpan' appears...").

[0040] The second analysis information for this example further includes the emotion vectors related to the representative frame and the text [anger: 22, expectation: 63, joy: 34, trust: 25, fear: 25, surprise: 56, sadness: 3, disgust: 3], and statistical values ​​calculated based on these vectors, such as emotion values ​​(normalized Euclidean norm of the vectors: 35.4), maximum emotion: 63 (expectation), and minimum emotion: 3 (sadness). This information also includes a correlation coefficient of "1.00" between the emotion values ​​and the moving average.

[0041] Because the example includes the second analysis information described above, the support device 1 can present a combination of representative frames, speech recognition results, emotion vectors, emotion values, maximum and minimum emotions, and generated advice for each scene of the video content. As a result, users can easily understand how to modify expressions to approach ideal emotional changes, while comparing the intensity and / or bias of emotions evoked in each specific scene with the content of the scene itself.

[0042] [Communication Unit 14] The specific configuration of the communication unit 14 is not particularly limited, as long as it is for connecting the support device 1 to the network N and performing communication. The communication unit 14 may be configured using, for example, a network card compatible with the Ethernet standard, a communication device compatible with wireless LAN, or other communication devices.

[0043] [Network N] The type of network N is not particularly limited as long as it is used for communication between the support device 1 and other devices. Examples of network N include the internet, a mobile phone network, and a wireless LAN.

[0044] [Terminal T] Terminal T is used by the user of support device 1. Terminal T performs processes such as receiving and / or displaying data transmitted from support device 1, transmitting user input to support device 1, and other processes. The types of terminal T may be, for example, portable terminals (e.g., laptop computers, smartphones, tablet computers) or stationary terminals (e.g., desktop computers).

[0045] [Main Flowchart of Support Processing] Figure 3 is a main flowchart showing an example of a preferred flow of support processing performed by the support device 1 of this embodiment. Figure 4 is a continuation from the previous figure. Figure 5 is a continuation from the previous figure. The following is a description of an example of a preferred flow of support processing performed by the support device 1, using Figures 3 to 5.

[0046] [Step S1: Determine whether to acquire content data] The control unit 11 performs a process to determine whether to acquire content data to be used for displaying content including multiple types of media using the content data acquisition unit 111 (content data acquisition determination step). If the control unit 11 determines to acquire the data, it moves the process to step S2; otherwise, it moves the process to step S3. The content data acquisition unit 111 is realized through the cooperation of the control unit 11, the storage unit 13, the communication unit 14, and other hardware components.

[0047] The above determination in this step is achieved by a procedure that determines whether to acquire the data if it falls under a predetermined case (e.g., a case in which a user's instruction regarding the upload of content data is received, or a case in which a user's instruction to start analysis, including the specification of content data, is received).

[0048] [Step S2: Acquire content data] The control unit 11 executes the process of acquiring content data using the content data acquisition unit 111 (content data acquisition execution step). The control unit 11 then moves the process to step S3.

[0049] The acquisition described above in this step is achieved, for example, by a procedure to receive content data transmitted by the user via terminal T, or by a procedure to acquire a resource specified using a Uniform Resource Identifier (URI) or other resource name identification means using the storage unit 13 or the communication unit 14.

[0050] When content analysis is ordered, content data corresponding to the content is specified, and the data is available, the support device 1 executes a series of processes to identify analysis units in line with the progression of the content. Steps S3 to S11 are an example of such processes.

[0051] [Step S3: Determine whether to identify analysis units based on the video] The control unit 11 uses the analysis unit identification unit 112 to perform a process to determine whether or not to identify analysis units based on the video (video-based identification determination step). If the control unit 11 determines that it should identify the units, it moves the process to step S5; otherwise, it moves the process to step S4. The analysis unit identification unit 112 is realized in cooperation with the control unit 11 and other hardware components. When an external device identifies the analysis units, the analysis unit identification unit 112 communicates with the external device using the communication unit 14.

[0052] The above determination in this step is achieved, for example, by a procedure that determines to identify the unit of analysis based on the video if it falls under one of the following predetermined cases (e.g., a case where the content data includes video media, a case where the content data includes video media and the primary medium of the content is specified as video media, or a case where the format of the content data matches one of the pre-specified ones).

[0053] Furthermore, if the support device 1 is configured and used as a support device for editing composite media content, primarily video (a support device for editing video media content), this step and the step of identifying analysis units using non-video media may be omitted, and a series of processes for identifying analysis units using video media may be performed instead.

[0054] [Step S4: Identifying analysis units based on non-video media] The control unit 11 uses the analysis unit identification unit 112 to perform a process to identify analysis units based on non-video media (non-video media-based identification step). The control unit 11 then moves the process to step S5.

[0055] An example of the procedure for achieving the above identification in this step is given below. Preferably, the procedure for achieving the above identification is configured so that the procedure is selected according to the type of content.

[0056] (Identifying units of analysis in manga content) Identifying units of analysis in manga or similar content is achieved, for example, by a series of steps that involve obtaining image media corresponding to each panel from the content data and identifying the sequence of these multiple image media in accordance with the progression of the content. In this case, each panel of the manga is identified as a unit of analysis. Since the units of analysis are the basic elements that make up the manga, it is possible to analyze changes in emotions in accordance with the progression at a fine level of detail.

[0057] If the image media included in the content data is pre-divided into frames, the analysis unit identification unit 112 acquires the frame-by-frame image media and identifies the order. The order identification may be, for example, based on the order specified when the content data was provided, or based on file names or other identification information.

[0058] If the image medium included in the content data contains multiple frames, the analysis unit identification unit 112 uses an image recognition model or other image recognition means to obtain the image medium divided into frames and identify the order. The order identification may be, for example, determined by the arrangement of the frames in the image medium containing multiple frames.

[0059] To suppress errors in frame-by-frame division, it is preferable that the image recognition means uses a pre-trained model that has been pre-trained using training data that associates image media containing multiple frames with frame-by-frame division.

[0060] Furthermore, in processing manga or similar content, it is preferable that the support device 1 acquires the text indicated by dialogue, monologues, and / or handwritten characters by performing character recognition processing on the image medium. This allows the support device 1 to perform emotional analysis including the text.

[0061] (Identifying the unit of analysis in text content) Identifying the unit of analysis in text or similar content is achieved, for example, by a series of steps that involve obtaining the text medium from content data and dividing the text medium based on the constituent units of the text. In this case, each divided text is identified as a unit of analysis. Since the unit of analysis is a block of elements that make up the text, it is possible to analyze changes in emotion along the progression with an appropriate level of granularity.

[0062] The division of a text medium based on its structural units is performed, for example, based on formal paragraphs or sentences. In the division of a text medium based on formal paragraphs, for example, each formal paragraph is identified as an analysis unit. In the division of a text medium based on sentences, for example, a predetermined number of sentences are identified as analysis units.

[0063] (Identifying units of analysis in meeting record content) Identifying units of analysis in meeting records or similar content is achieved, for example, by a series of steps that involve obtaining the medium containing the spoken content (e.g., audio medium, written medium) from the content data, preparing it for processing as written medium as necessary, and then dividing the written medium based on the number of sentences. In this case, for example, a predetermined number of sentences are identified as units of analysis.

[0064] Meeting records contain statements made by multiple speakers, either consecutively or alternately. Each statement is typically represented as a single sentence in the written medium. However, written materials, particularly those obtained from audio recordings of meeting records using speech recognition, may not be properly divided into paragraphs. Procedures using segmentation based on the number of sentences can identify appropriate units of analysis even in such written materials.

[0065] [Step S5: Video Analysis Unit Identification Loop] The control unit 11 uses the analysis unit identification unit 112 to execute the process of starting the video analysis unit identification loop (video analysis unit identification loop start step). The control unit 11 then moves the process to step S6.

[0066] In order to identify analysis units for a wide range of videos by using multiple analysis unit identification criteria, it is preferable that the analysis unit identification unit 112 in the video analysis unit identification loop identifies analysis units using at least one of the magnitude of changes in the video included in the content data and the video length.

[0067] In the video analysis unit identification loop, it is preferable that the analysis unit identification unit 112 is configured to identify the analysis unit using at least a first means for identifying the analysis unit based on the magnitude of changes in the video, and a second means for identifying the analysis unit if the first means fails to identify the analysis unit. This allows the support device 1 to solve the problem that the identification of the analysis unit based on the magnitude of changes in the video is not always successful through fallback processing.

[0068] Steps S6 to S11 are an example of a video analysis unit identification loop that uses a first means for identifying an analysis unit based on the magnitude of changes in the video, a second means for identifying an analysis unit if the first means fails to identify the analysis unit, and a third means for identifying an analysis unit using the video length if the second means fails to identify the analysis unit.

[0069] [Step S6: Attempt to identify by first means] The control unit 11 uses the analysis unit identification unit 112 to perform a process to attempt to identify the analysis unit by first means, which identifies the analysis unit based on the magnitude of changes in the video (first means attempt step). The control unit 11 then moves the process to step S7.

[0070] (Regarding the first means) The first means may be a means implemented by the support device 1 alone, a means implemented by the cooperation of the support device 1 and an external device, or a means implemented by an external device. The first means may be, for example, a means of dividing a video into scenes based on the magnitude of changes in the video. Alternatively, the first means may be a means of dividing a video into scenes based on both the magnitude of changes in the video and an audio medium played in sync with the video.

[0071] Examples of the first method include a method for dividing the analysis unit when the histogram difference between adjacent frames exceeds a threshold, a method for dividing the analysis unit when the difference in edge information between adjacent frames exceeds a threshold, a method for dividing the analysis unit when the structural similarity falls below a threshold, or a method combining several of these.

[0072] Furthermore, as an example of a first means for dividing a video into scenes based on both the magnitude of changes in the video and the audio medium played in sync with the video, in addition to identifying analysis units based on the magnitude of changes in the video, a means for dividing analysis units when the length of a silent interval exceeds a threshold or when the change in acoustic characteristics exceeds a threshold can be mentioned.

[0073] To enable fallback processing to the second means, the first means is preferably configured to notify of failure if an attempt to identify the unit of analysis fails. For example, the first means executes a process if the processing time required for identification exceeds a threshold, treating it as an attempt failure. This prevents an increase in the processing time in the first means.

[0074] In order to reduce the expected processing time required to identify the analysis unit, in the video analysis unit identification loop of this example, it is preferable that the first means is a means that takes less processing time to identify the unit than the second means described later. For this reason, it is preferable that the threshold for processing time in the first means is set to be shorter than that of the second means.

[0075] To suppress errors in identifying the analysis unit, the first method is to use a pre-trained model that has been pre-trained using training data that associates video media with division into analysis units.

[0076] [Step S7: Determining whether identification by the first means failed] The control unit 11 uses the analysis unit identification unit 112 to perform a process to determine whether identification by the first means failed (first means failure determination step). If the control unit 11 determines that it failed, it moves the process to step S8; otherwise, it moves the process to step S11.

[0077] [Step S8: Attempting identification by second means] The control unit 11 uses the analysis unit identification unit 112 to perform a process to attempt to identify the analysis unit by a second means that identifies the analysis unit based on the magnitude of changes in the video (second means trial step). The control unit 11 then moves the process to step S9.

[0078] (Regarding the second method) In principle, the second method may be the same as the first method. However, in order to avoid trial failures due to the same reasons as the first method, it is preferable that the second method is a method in which the method and / or threshold setting differs from that of the first method. If only the threshold setting differs from that of the first method, it is preferable that the threshold for the second method is set so that the probability of trial failure is lower than that of the first method.

[0079] In order to perform fallback processing to the third means, the second means is preferably configured to notify of failure if an attempt to identify the unit of analysis fails. For example, the second means executes processing as an attempt failure if the processing time required for identification exceeds a threshold. This prevents an increase in the processing time in the second means. The relationship between the setting of the processing time threshold in the second means and the first means is as described in the section on the first means.

[0080] [Step S9: Determining whether identification by the second means failed] The control unit 11 uses the analysis unit identification unit 112 to perform a process to determine whether identification by the second means failed (second means failure determination step). If the control unit 11 determines that it failed, it moves the process to step S10; otherwise, it moves the process to step S11.

[0081] [Step S10: Identification by fixed-length division] The control unit 11 uses the analysis unit identification unit 112 to perform a process to identify analysis units using a third means that identifies analysis units based on the video length (third means execution step). The control unit 11 then moves the process to step S11. The third means is, for example, a means that divides the video so that the video length related to the analysis unit is a fixed length.

[0082] In the example of the video analysis unit identification loop described above, the analysis units are identified in the following order: a first means that identifies analysis units by capturing only relatively large changes with a short processing time; a second means that identifies analysis units by capturing relatively small changes as well, although the processing time is longer; and a third means that identifies analysis units regardless of changes, although the processing time is shortest. As a result, the analysis unit identification unit 112 is expected to identify analysis units that are able to prevent the omission of changes in emotional movement related to the video medium and keep the number of analysis units to a minimum within a limited processing time, without unnecessarily increasing the processing time required for identifying analysis units.

[0083] [Step S11: Terminate if there are no remaining video analysis targets] The control unit 11 uses the analysis unit identification unit 112 to perform a process to determine whether or not there are any remaining video analysis targets (video analysis unit identification loop termination determination step). If the control unit 11 determines that there are no remaining targets, it terminates the video analysis unit identification loop and moves the process to step S12; otherwise, it continues the video analysis unit identification loop and moves the process to step S5. The remaining video analysis targets referred to here are the remaining parts of the video that are not included in the analysis units already identified by the loop described above.

[0084] After the analysis unit is identified, the support device 1 uses the emotion information acquisition unit 113 to perform a series of processes that cause the converter to output emotion information for each analysis unit, including the intensity of multiple types of emotions evoked by a part of the content related to that analysis unit. Through this series of processes, the emotion information acquisition unit 113 can acquire the outputted emotion information. In this process, the emotion information acquisition unit 113 inputs data corresponding to the analysis unit based on the content data for each analysis unit into the converter, causing the converter to output the aforementioned emotion information. Steps S12 to S16 are an example of this series of processes.

[0085] [Step S12: Emotion Index Calculation Loop] The control unit 11 executes a process to start the emotion index calculation loop using the emotion information acquisition unit 113 (emotion index calculation loop start step). The control unit 11 then moves the process to step S13. The emotion information acquisition unit 113 is realized through the cooperation of the control unit 11, the storage unit 13, and other hardware components. When calculating the emotion index in an external device, the emotion information acquisition unit 113 communicates with the external device using the communication unit 14.

[0086] In the emotion index calculation loop of this example, the emotion information acquisition unit 113 executes the processes from step S13 to step S15 for each unprocessed analysis unit.

[0087] [Step S13: Generate data to be calculated] The control unit 11 uses the emotion information acquisition unit 113 to perform the process of generating data to be calculated (calculation data generation step). The control unit 11 then moves the process to step S14.

[0088] The data to be calculated in this step is data corresponding to the analysis unit based on the content data described above, and is generated, for example, by a procedure that extracts the portion corresponding to the analysis unit from the data of each medium included in the content data.

[0089] Incidentally, calculating and acquiring emotional information based on video media usually requires more processing time than calculating and acquiring it based on image media. Therefore, in order to reduce the processing time required to acquire emotional information, it is preferable that the extraction in video media is generated by a procedure that extracts representative images (representative frames) of scenes corresponding to the analysis unit. When calculating emotional information based on emotions indicated by movement, it is preferable that the extraction in video media includes a procedure that extracts a part of the video as a representative clip.

[0090] [Step S14: Acquire emotion information] The control unit 11 uses the emotion information acquisition unit 113 to input the calculation target data corresponding to the above-mentioned analysis unit into the converter, and causes the converter to output emotion information including the intensity of multiple types of emotions evoked by a part of the content related to the analysis unit, thereby executing the process of acquiring the emotion information (emotion information acquisition step). The control unit 11 then moves the process to step S15.

[0091] In this step, the "converter" is configured to analyze the input data to be calculated and output emotional information that includes the intensity of multiple types of emotions evoked by the content represented by the data. "Multiple types of emotions" are, for example, the eight basic emotions in Plutchik's Wheel of Emotions. The eight basic emotions in Plutchik's Wheel of Emotions are anger, expectation, joy, trust, fear, surprise, sadness, and disgust.

[0092] Emotional information is not particularly limited as long as it includes the intensity of multiple types of emotions evoked by the content, and for example, it includes an emotion vector whose elements are the intensity of multiple types of emotions evoked by the content. Hereinafter, "the intensity of multiple types of emotions evoked by the content as indicated by the content data" will also be simply referred to as "the emotion vector corresponding to the content data."

[0093] The converter is configured using a pre-trained model obtained by training data that associates content data with emotion vectors corresponding to the content data. This model includes, for example, a language model, an image recognition model, and a multimodal model. In order to obtain emotion information including the changes in emotion evoked by the combination, it is preferable that the converter be configured using a multimodal model.

[0094] In order to output the emotions evoked by sound effects, music, and other non-verbal content in the audio medium contained in the content data, the converter may be configured to analyze the data relating to the audio medium itself and output emotional information including the intensity of multiple types of emotions evoked by the content indicated by the data. In order to perform analysis that is consistent with emotional information relating to text medium, the converter may be configured to analyze the data obtained by converting the audio medium into text using speech recognition means and output the emotional information.

[0095] Preferably, the converter is configured to output emotional information, including an emotional vector corresponding to the content data, when data relating to multiple media included in the content data is input. This allows the converter to analyze not only the emotional changes evoked by each medium, but also the emotional changes evoked by combining each medium in a multimedia content.

[0096] In order to pre-train the emotional changes evoked by combinations, it is preferable that the model relating to the converter is pre-trained using training data that associates data relating to multiple media with corresponding emotional vectors. Furthermore, the converter may be configured to generate emotional information including emotional vectors corresponding to each of the multiple media using pre-trained models corresponding to each media, and to output emotional information including integrated emotional vectors corresponding to the multiple media by integrating the emotional vectors corresponding to the multiple media element by element.

[0097] The converter may be implemented within the support device 1, or it may be implemented in an external device separate from the support device 1 included in system S. By implementing the converter in an external device, it is possible to shorten the processing time through parallel processing between the support device 1 and the external device.

[0098] In identifying emotions based on representative images extracted from video media or other image media, it is preferable that the converter be configured to output regions in the image that serve as the basis for identifying the emotions, in order to provide the basis for the identified emotions. The format of the output is not particularly limited. Examples of such formats include a format of digital data that associates data indicating the location and shape of the region with the type, intensity, and / or contribution to identification of the emotion represented by the region; a format of image data that expresses the contribution to identification in the region by shades of gray; a format of image data that expresses the intensity of the emotion represented by the region by shades of gray; and a format of image data that expresses the type of emotion represented by the region by color.

[0099] In identifying emotions based on textual media, it is preferable that the converter be configured to output elements in the textual media (e.g., words, phrases) that served as the basis for identifying the emotion, in order to provide the basis for the identified emotion. The format of the output is not particularly limited. Examples of such formats include a format of digital data that associates the element with the degree of contribution of the domain to emotion identification, and a format of digital data that associates the element with the emotion represented by the domain.

[0100] [Step S15: Calculate emotion index based on emotion information] The control unit 11 uses the emotion information acquisition unit 113 to perform the process of calculating an emotion index based on the emotion information described above (emotion index calculation step). The control unit 11 then moves the process to step S16.

[0101] The calculation of the emotion index in this step is achieved, for example, by mathematical and / or statistical processing of the emotion information. The emotion information acquisition unit 113 calculates, for example, the norm of the emotion vector or other indices indicating its intensity as an emotion index. The method for calculating the intensity of the emotion vector is not particularly limited, and may be, for example, a method in which one of the maximum value norm, Euclidean norm, p-order mean norm, or norm defined by the inner product is calculated and the said norm is used as the intensity.

[0102] Furthermore, in order to reduce the computational load in the analysis based on the complex emotions evoked by the content and to obtain useful indicators, it is preferable that the calculation of the emotion indicator in this step includes a procedure for calculating a statistical value that treats each element included in the emotion vector as a population. For the reasons mentioned above, it is particularly preferable that the statistical value includes both the median (emotion value) and the mean (emotion mean), which are a set of statistical values ​​that require particularly little computation. The procedure for calculating the median when the number of elements included in the emotion vector is even may be, for example, a procedure for calculating the average of the two elements that are closest in rank to the center.

[0103] To perform an analysis based on the range of emotions evoked by the content, the calculation of the emotion index in this step preferably includes a procedure for calculating the maximum and minimum values ​​of each element of the emotion vector as part of the emotion index. To perform an analysis that includes emotions that are particularly strongly evoked, the emotion index preferably further includes the type of emotion corresponding to the maximum value. To perform an analysis that includes emotions that are not evoked as strongly as other emotions, the emotion index preferably further includes the type of emotion corresponding to the minimum value. Furthermore, for an analysis using the range of emotions based on a statistical index, the calculation preferably includes a procedure for calculating the standard deviation of each element of the emotion vector as a population as part of the emotion index. To perform an analysis based on specific emotions evoked by the content, the emotion index in this step preferably includes each element of the emotion vector as part of the emotion index.

[0104] [Step S16: Terminate if no analysis units remain] The control unit 11 uses the emotion information acquisition unit 113 to determine whether or not there are any unprocessed analysis units remaining (emotion index calculation loop termination determination step). If the control unit 11 determines that there are no remaining units, it terminates the emotion index calculation loop and moves the process to step S17; otherwise, it continues the emotion index calculation loop and moves the process to step S12.

[0105] After the emotional indicators are identified, the support device 1 uses an analysis information generation unit 114 to perform a series of processes to generate analysis information related to the editing of the content, based on the changes in the emotional indicators calculated based on the emotional information in line with the progression of the content and on the criteria related to the emotional indicators. Through this series of processes, the analysis information generation unit 114 can acquire the generated analysis information.

[0106] In this process, it is preferable that the analysis information generation unit 114 includes a procedure for inputting the data showing the above-mentioned changes into a large-scale language model to generate text related to the analysis information. This allows the analysis information generation unit 114 to generate text based on the large-scale language model that includes not only the emotions that each medium constituting the content evokes in the viewer, but also descriptions of the emotions that are evoked in the viewer due to the combination of each medium. Steps S17 to S18 are an example of this series of processes.

[0107] [Step S17: Calculate changes in the emotional index in line with its progression] The control unit 11 uses the analysis information generation unit 114 to perform a process to calculate changes in the emotional index in line with its progression (change calculation step). The control unit 11 then moves the process to step S18. The analysis information generation unit 114 is realized through the cooperation of the control unit 11 and the storage unit 13 and other hardware components. When calculating the changes in an external device, the analysis information generation unit 114 communicates with the external device using the communication unit 14. Changes in the emotional index in line with its progression are indicated, for example, by one or more analysis values ​​or similar analysis results.

[0108] (Correlation coefficient with moving average) The calculation of the change in this step preferably includes a series of steps to calculate the moving average of one of the indicators included in the sentiment indicator in line with the progression of the content, and then calculate the correlation coefficient between the indicator and the moving average.

[0109] The moving average of the sentiment index corresponds to the overall strength of emotion in the surrounding analysis units along the progression of the content, not just the corresponding analysis unit. Therefore, the correlation coefficient will be large in areas where the index does not change much and the index itself is small, and in areas where the index changes rapidly and the index itself is large. On the other hand, the correlation coefficient will be small in areas where the index does not change much but the index itself is large, and in areas where the index changes rapidly but the index itself is small. Thus, by calculating the change in the sentiment index along the progression, including the series of steps described above, it is possible to evaluate which analysis unit corresponds to the peaks in the content where the index changes rapidly and becomes stronger.

[0110] The moving average may be any type of moving average, such as the simple moving average, weighted moving average, or exponential moving average. As an example of a moving average that emphasizes the unit of analysis under focus compared to surrounding units of analysis and reduces processing time, a weighted moving average can be used, in which the indicator corresponding to the unit of analysis under focus is multiplied by a weight of 1 / 2, and the indicators corresponding to the preceding and succeeding units of analysis are multiplied by a weight of 1 / 4, and the sum of these is calculated.

[0111] To reduce processing time, it is preferable for the analysis information generation unit 114 to calculate the correlation coefficient between the indicator corresponding to the analysis unit under focus and the preceding and succeeding analysis units, and the corresponding number of moving averages. As a correlation coefficient, for example, the Pearson correlation can be used, as it can measure linear relationships while reducing processing time.

[0112] (Regarding other changes) The calculation of such changes in this step may further include, for example, steps to calculate the mean of the indicator, the maximum and minimum values ​​of the indicator, the mean range of the indicator, the rate of change of the indicator, and the mean start-end difference of the indicator.

[0113] The average range of an indicator is the range of change in the indicator within a given range. The rate of change of an indicator is the percentage increase or decrease in the indicator within a given range. The average start-end difference of an indicator is the amount by which the indicator changed within a given range. The start-end difference of an indicator is the amount by which the indicator changed within a given range. The average of an indicator is the average of the indicator within a given range. The given range is not particularly limited and may be, for example, the entire text, a range corresponding to multiple consecutive units of analysis, or a range specified by the user.

[0114] By including a procedure for calculating the average width or rate of change in this step, the analysis information generation unit 114 can analyze the pacing and emphasis in the content. By including a procedure for calculating the average start-to-end difference and / or start-to-end difference in this calculation, the analysis information generation unit 114 can analyze whether the analysis unit is a complete part of the content, a continuing part, or a part that is neither of these. By including a procedure for calculating the average in this calculation, the analysis information generation unit 114 can analyze the emotions evoked by the content.

[0115] (Specifying indicators) When the emotional indicator includes multiple indicators, the analysis information generation unit 114 preferably includes a procedure for calculating the change in one indicator specified by the user in accordance with the progression.

[0116] To ensure that changes in the overall intensity of emotions are not overlooked in the analysis, it is preferable that the analysis information generation unit 114 includes a procedure for calculating changes in emotional values ​​in line with their progression, regardless of whether it is specified or not. To ensure that calculation targets that have a significant effect relative to the short processing time are included in the analysis, it is preferable that the analysis information generation unit 114 includes a procedure for calculating the average of the indicators, regardless of whether it is specified or not.

[0117] [Step S18: Generate analysis information] The control unit 11 uses the analysis information generation unit 114 to perform a process to generate analysis information related to content editing based on the changes in the emotional indicators in line with their progression and the criteria related to the emotional indicators (analysis information generation step). The control unit 11 then moves the process to step S19.

[0118] Preferably, the analytical information generated in this step includes text describing the characteristics of the content, generated based on the changes and criteria, so that users can edit based on objectively analyzed features. Preferably, the analytical information generated in this step includes text providing editing advice, generated based on the changes and criteria, so that users can edit with reference to advice based on objective analysis.

[0119] The generation of analysis information in this step includes, for example, a procedure to obtain an analysis result for each analysis value showing the change, associated with a given range to which the analysis value corresponds. In this procedure, for example, for the correlation coefficient, a correspondence table stored in the memory unit 13 is referred to (e.g., a table where a value of 0 or more but less than 0.25 corresponds to the first analysis result, 0.25 or more but less than 0.5 corresponds to the second analysis result, 0.5 or more but less than 0.75 corresponds to the third analysis result, and 0.75 or more corresponds to the fourth analysis result), and the second analysis result is obtained when the correlation coefficient is 0.33.

[0120] (Regarding agreement with a given pattern) The generation of analysis information in this step preferably includes a procedure to identify the pattern with the highest agreement rate with the changes in the analysis value as it progresses from among a plurality of patterns stored in the storage unit 13 in advance, and to generate analysis information based on that pattern and the agreement rate.

[0121] As a result, the analysis information generation unit 114 can identify patterns with a high degree of agreement with the content from among various patterns exemplified by, for example, the Man in a Hole type, the Success type, the Tragedy type, the Comeback type, the Downfall type, the Cinderella type, and the Oedipus type, and generate analysis information that includes advice on how to approach the ideal change corresponding to that pattern. Furthermore, the analysis information generation unit 114 can use this agreement rate as an overall score for the content.

[0122] (Regarding generation using a large-scale language model) The generation of analytical information in this step preferably includes a procedure for inputting the data showing the above-mentioned changes into a large-scale language model to generate text related to the analytical information. This allows the analytical information generation unit 114 to generate text based on the large-scale language model that includes not only the emotions that each medium constituting the content evokes in the viewer, but also descriptions of the emotions that are evoked in the viewer due to the combination of each medium.

[0123] In order to generate text based on the criteria for the change, the procedure for inputting the data showing the change described above into a large-scale language model to generate text related to the analysis information preferably involves further inputting text showing the criteria.

[0124] To generate text based on insights into emotional changes not included in general corpora, the large-scale language model is preferably a pre-trained model that has been trained using training data that indicates criteria for such changes. The training data that indicates criteria for such changes is, for example, training data that associates data indicating such changes with corresponding texts.

[0125] In this case, it is preferable that the learning data includes data that associates the changes with text that describes the characteristics of the content, so that users can edit based on objectively analyzed features. Furthermore, it is preferable that the learning data includes data that associates the changes with text that provides advice on editing, so that users can edit by referring to advice based on objective analysis.

[0126] In order to generate text that includes descriptions based on a deeper analysis of the emotions evoked in viewers due to the combination of each medium, it is preferable that the large-scale language model is a pre-trained model obtained by pre-training using training data that associates data showing the changes in multiple mediums with corresponding texts.

[0127] [Step S19: Command to display analysis information] The control unit 11 generates visual representation data for displaying the analysis information mentioned above using the visual representation data generation unit 115, and executes a process to command a display based on the data (display command step). The control unit 11 returns the process to step S1 and repeats the process from step S1 to step S19.

[0128] The visual representation data generation unit 115 is realized through the cooperation of the control unit 11, the storage unit 13, and other hardware components. When generating and / or displaying the data in cooperation with an external device, the visual representation data generation unit 115 communicates with the external device using the communication unit 14.

[0129] The following is an example of a visual representation included in the visual representation data generated in this step.

[0130] (Graphic representation of emotional indicators and their changes) In order to make it easier for users to visually recognize emotional indicators and / or their changes, it is preferable that the data for visual representation be a graph representation of emotional indicators and / or their changes.

[0131] To facilitate the recognition of changes in the overall intensity of emotions, the graph representation preferably includes line graphs, curve graphs, bar graphs, and other graph representations related to the emotion values. To facilitate the recognition of changes in the correlation coefficient related to the emotion values, the graph representation preferably includes line graphs, curve graphs, bar graphs, and other graph representations related to the correlation coefficient. To facilitate the recognition of changes in the range of emotion intensity for each analysis unit, the graph representation preferably includes graph representations showing the maximum and minimum values ​​of the elements of the emotion vector (e.g., a range bar graph showing the range from the minimum to the maximum value).

[0132] To make it easier for users to recognize changes in the indicators they specify, the graph representation preferably includes line graphs, curve graphs, bar graphs, and other graph representations related to the indicators. Examples of the specified indicators include elements specified by the user from each element of the emotion vector (e.g., elements indicating the intensity of emotion corresponding to the emotion of anger specified by the user).

[0133] (Representations indicating content features) In order for users to recognize the features of the content, it is preferable that the visual representation data includes representations indicating the features of the content. Such representations include, for example, one or more of the following: text included in the analysis information, a graph representation, an icon representation and / or text indicating the size of an indicator indicating a feature, and a graph representation, an icon representation and / or text indicating the ratio of the size of the indicator indicating a feature to a standard.

[0134] Characteristics of the content relating to the expression include, for example, characteristics shown by the pattern described above, characteristics shown by the degree of agreement with the pattern described above, characteristics shown by the average emotional value, characteristics shown by the rate of change in emotional value, and characteristics shown by the difference between the starting and ending emotional values.

[0135] (Expressions indicating advice) To convey advice regarding content editing to users, the visual representation data preferably includes expressions indicating such advice. Such expressions include, for example, text included in the analytical information. Such expressions may be expressed in whole or in part together with some or all of the expressions that describe the characteristics of the content, or they may be expressed separately. To aid visual recognition by users, such expressions preferably include illustrations, icons, background images, or other pictorial representations that correspond to the size or content of the indicators included in the analytical information.

[0136] (Expressions indicating the basis for emotion identification) In order to convey the basis for emotion identification to the user, it is preferable that the visual representation data includes expressions indicating said basis. In representative images extracted from video media or other image media, the expression may be represented, for example, by superimposing an area in the image that particularly contributed to emotion identification with a shade indicating the intensity of the corresponding emotion, or by superimposing an area in the image that indicates the type of corresponding emotion. In text media, the expression may be represented, for example, by a visual representation that associates elements in the text media (e.g., words, phrases) with the emotions corresponding to those elements.

[0137] [Effects of support processing] As described above, the support device 1 of this embodiment includes a content data acquisition unit 111 which acquires content data to be used for displaying content including multiple types of media (steps S1 to S2), an analysis unit identification unit 112 which identifies analysis units in line with the progression of the content based on this content data (steps S3 to S11), an emotion information acquisition unit 113 which inputs data corresponding to the analysis unit based on this content data to a converter for each analysis unit and outputs emotion information including the intensity of multiple types of emotions evoked by a part of the content related to the analysis unit (steps S12 to S16), and an analysis information generation unit 114 which generates analysis information related to content editing based on the changes in the emotion index calculated based on this emotion information in line with the progression and criteria related to the emotion index, inputs the data showing the changes to a large-scale language model to generate text related to the analysis information (steps S17 to S18), and executes support processing.

[0138] In this support process, the emotion information acquisition unit 113 inputs data corresponding to analysis units in line with the progression of the content, which have been identified based on the content data, into the converter, and outputs emotion information including the intensity of multiple types of emotions evoked by a part of the content related to the analysis unit (steps S12 to S16). Then, the analysis information generation unit 114 generates analysis information related to the editing of the content based on the changes in the emotion index calculated based on the emotion information in line with the progression and the criteria related to the emotion index (steps S17 to S18).

[0139] At this time, the analysis information generation unit 114 inputs data showing changes in line with the progression of the emotion index into a large-scale language model to generate text related to the analysis information (step S18). The large-scale language model typically generates text based on trained parameters that represent the contextual relationships of the input token sequence. Therefore, the invention can generate text based on the large-scale language model that includes not only the emotions evoked in the viewer by each medium constituting the content, but also descriptions of the emotions evoked in the viewer due to the combination of each medium.

[0140] As a result, the support device 1 that performs the above-mentioned support processing can provide a suggestion function for editing multimedia content based on multimodal sentiment analysis of the content. Therefore, one aspect of the present invention relating to the support device 1 can provide a technical means for supporting the editing of multimedia content that takes into account changes in line with the progression of emotions evoked in the viewer.

[0141] The above-described support processing may be configured such that the analysis unit identification unit 112 identifies the analysis unit using at least one of the magnitude of change in the video included in the content data and the video length (steps S5 to S11).

[0142] As a result, the support device 1 can identify analysis units with appropriate granularity according to the video, and can appropriately analyze changes in line with the progression of emotions in content including video. Consequently, it can capture both rapid scene changes and gradual developments in video media without excess or deficiency, thereby improving the accuracy of content analysis.

[0143] Furthermore, the above-described support processing can be configured such that the analysis unit identification unit 112 identifies the analysis unit using at least a first means for identifying the analysis unit based on the magnitude of changes in the video, and a second means for use when the first means fails to identify the analysis unit (steps S6 to S9).

[0144] As a result, the support device 1 can identify the analysis unit through fallback processing even in cases where the first means fails. Consequently, it is possible to stably obtain the analysis unit according to the characteristics or processing status of the video, thereby improving the reliability of sentiment analysis on content including video.

[0145] Furthermore, the above-described support processing may be configured such that the analysis unit identification unit 112 further uses a third means, in which a first means identifies an analysis unit based on the magnitude of changes in the video, and if the identification by the first means fails, a second means identifies an analysis unit based on the magnitude of changes in the video, and if the identification by the second means fails, a third means identifies an analysis unit based on the video length (steps S6 to S10).

[0146] As a result, even if neither the first nor the second means can identify an analysis unit based on the magnitude of changes in the video, the support device 1 can ultimately identify a fixed-length analysis unit based on the video length. Consequently, an analysis unit can be reliably obtained in all cases, enabling continuous and uninterrupted sentiment analysis of content including videos.

[0147] In addition, the above-mentioned support processing can be configured to identify analysis units by executing in the following order: a first means that identifies analysis units by capturing only relatively large changes with a short processing time; a second means that identifies analysis units by capturing relatively small changes as well, although the processing time is longer; and a third means that identifies analysis units regardless of changes, although the processing time is shortest (steps S6 to S10). As a result, the analysis unit identification unit 112 is expected to identify analysis units that are as suitable as possible for the content of the video medium within a limited processing time, without unnecessarily increasing the processing time required for identifying analysis units.

[0148] Therefore, the present invention provides a technical means to support the editing of composite media content, which primarily consists of video, by performing analysis that reduces processing time while maintaining an appropriate level of granularity that takes into account changes in line with the progression of emotions evoked in the viewer.

[0149] <Example of Use> The following is an example of using the support device 1 of this embodiment.

[0150] [Initialization of Support Device 1] The user accesses Support Device 1 from terminal T and starts the program that executes the support processing of this embodiment. As an initialization step, the program loads the analysis model, moving average window settings, correlation coefficient calculation method, and other default values, and displays the analysis screen.

[0151] [Specifying Video Content] The user selects "New Analysis > Content Including Video" and specifies the video file to be analyzed (e.g., MP4). Support device 1 acquires the video (from step S1 to step S2).

[0152] [Automatic identification of analysis units (scenes)] The support device 1 recognizes that the content includes video media and identifies analysis units (scenes) based on the video (steps S3 to S11).

[0153] [Content Analysis] The support device 1 extracts a representative frame or short representative clip for each analysis unit and inputs it into the converter to obtain an emotion vector whose elements are the intensity of Plutchik's eight basic emotions (steps S12 to S14). Then, the support device 1 calculates emotion values ​​(e.g., normalized Euclidean norm), maximum / minimum emotion, their standard deviation, correlation coefficient with moving average, and other emotion indicators from the emotion vector of each scene.

[0154] [Generation of Analysis Information] The support device 1 obtains analysis values ​​(e.g., mean, rate of change, difference from start to finish) based on changes in the emotion index as it progresses, inputs them into a large-scale language model along with the reference text, and generates analysis information text that includes content features, summarization, and editing advice (steps S17 to S18).

[0155] [Generation and display of visual representations] The visual representation data generation unit 115 generates visual representation data related to the analysis information and uses this data to display a visual representation of the analysis information on the terminal T (step S19).

[0156] [Example of Visual Representation] Figure 6 is an example of a display related to analytical information. This example includes the following visual representations.

[0157] (1) Header: In the header section at the top of Figure 6, directly below "Work Analysis," a visual summary of the content is displayed. To the left of the header section is the total score (60 points) and a meter-like diagram showing the percentage relative to the maximum score. To the right of the header section is the final evaluation (rank C) and an AI summary generated by a large-scale language model ("To get closer to the 'man in a hole' type, it is important to first set a high baseline for the protagonist at the beginning of the story, and then to deeply depict the emotional 'valley.' And then the frustration...").

[0158] (2) Graph: In the upper center of Figure 6, directly below the header, a graph is displayed showing the changes in the emotion value and emotion vector as they progress. This graph superimposes a solid line graph of the emotion value, a dotted line graph of the correlation coefficient, and a range bar graph showing the range from the minimum to the maximum value of each element of the emotion vector in each analysis unit. When data is selected by mouseover or other means, the tooltip displays "emotion value," "average emotion," "ideal emotion value," and other detailed data.

[0159] (3) Advice: In the lower center of Figure 6, directly below the graph, the text of the speech recognition results of the audio contained in the video content ("Around the time of the Meiji Restoration, those who liked Western culture were sometimes mocked as 'Westernized,' but those who couldn't keep up with the times were said to be 'Tenpo old men'...") and representative frames of the scene (from left to right: "A frame depicting a man surprised by Western culture," "A frame depicting a man kneading dough," "A frame depicting a man with his arms crossed glaring at the dough," "A frame depicting the entrance to a liquor store") are displayed along with editing advice generated by a large-scale language model ("Let's make the characters' expressions and reactions richer to emphasize their surprise and bewilderment at Western culture. Let's use backgrounds and sound effects to mask the uncertainty of the rumors and encourage the reader to empathize. 'Anpan' appears...") is displayed as a "suggestion."

[0160] (4) Indicators: Below the advice, in the lower part of Figure 6, emotional indicators and analytical values ​​showing change are listed along with their satisfaction rates relative to the criteria. In the lower part of Figure 6, from left to right, "Graph Correlation" (0.33, satisfaction rate 30%), "Average Emotion" (51.2, satisfaction rate 100%), and "Emotion Change Rate" (0.58, satisfaction rate 50%) are presented in card format.

[0161] Figure 7 shows an example of a detailed display of video analysis information. In the upper center and lower left of this example, representative images showing people's faces are displayed. In the lower right of this example, there is a visual representation indicating the contribution to emotion identification, where areas with a high contribution to emotion identification in the representative image are shown in red (bright rectangular areas in the lower right image), and areas with a low contribution to emotion identification are shown in blue (dark areas in the lower right image). This allows users to consider editing the video to make the emotions more desirable based on this visual representation.

[0162] Figure 8 is an example of a detailed display of text analysis information. In the center of the example, a sentence that is part of the text medium related to the analysis unit is shown: "Foemina-sama, no matter how much you rush, the cake won't run away." Below the example, there is a visual representation indicating the contribution of each element to emotion identification, with the sentence elements such as "run away," "won't run away," "cake," "sama," "you," "Foemina," "no matter how much you rush," "so much," and "rush" arranged in order from the least to the most important to identify the emotion. This allows users to consider editing the text to make the emotion more desirable based on this visual representation.

[0163] [Utilization of Analysis Results] Users can grasp an overview of the emotional dynamics in the content being edited through the overall summary and graphs. Users can gain a more detailed understanding of the emotional dynamics in the content through the detailed data shown in the tooltip that appears when the mouse hovers over the graph, and the analysis data for a specified range that appears when a range is selected.

[0164] Furthermore, users can refer to advice and indicators, or instruct the support device 1 to display an ideal emotion curve in the graph area, and edit the content to bring the changes in the progression of emotions evoked by the content closer to an ideal curve. Users can then have the support device 1 re-analyze the edited content to confirm whether the editing direction is appropriate, thereby efficiently improving the content.

[0165] <Note> Within the scope of the concept of the present invention, a person skilled in the art can conceive of various modifications and alterations, and it is understood that such modifications and alterations also fall within the scope of the present invention. For example, any addition, deletion, or design change of components, or addition of processes or changes in conditions, made by a person skilled in the art to the above-described embodiment, is also included within the scope of the present invention, as long as it retains the gist of the present invention. [Explanation of symbols]

[0166] S System 1 Support equipment 11 Control Unit 111 Content Data Acquisition Unit 112 Identification of Analysis Units 113 Emotion information acquisition unit 114 Analysis information generation section 115 Visual representation data generation unit 13 Storage section 131 Analysis Information Database 14 Communications Department N Network T terminal

Claims

1. A content data acquisition unit that acquires content data to be used for displaying content including multiple types of media, An analysis unit identification unit identifies an analysis unit in accordance with the progression of the content based on the content data, For each analysis unit, an emotion information acquisition unit inputs data corresponding to the analysis unit based on the content data into a converter and outputs emotion information including the intensity of multiple types of emotions evoked by a part of the content related to the analysis unit. An analysis information generation unit generates analysis information related to the editing of the content based on the changes in the emotion index calculated based on the aforementioned emotion information in line with the aforementioned progression and criteria related to the said emotion index, and inputs the analysis values ​​into a large-scale language model to generate text related to the said analysis information. Equipped with, The aforementioned analytical values ​​show changes in line with the progression described above. A support device for editing multimedia content.

2. The support device according to claim 1, wherein the analysis unit identification unit is configured to identify the analysis unit using at least one of the magnitude of change in the video and the video length included in the content data.

3. The support device according to claim 2, wherein the analysis unit identification unit is configured to identify the analysis unit using at least a first means for identifying the analysis unit based on the magnitude of the change, and a second means for identifying the analysis unit when the first means fails to identify the analysis unit.

4. The support device according to claim 3, wherein the analysis unit identification unit is configured such that the second means identifies the analysis unit based on the magnitude of the change, and if the second means fails to identify the analysis unit, the third means further identifies the analysis unit based on the video length.

5. The support device according to claim 1, wherein the analysis values ​​include at least one of graph correlation, average emotion, and emotion change rate.