Content Similarity Calculation System and Content Search / Recommendation System
The content similarity calculation system addresses the challenge of capturing temporal changes in media content meaning and impression by extracting semantic waveforms and calculating cosine similarity in the frequency domain, thereby enhancing content search and recommendation systems.
Patent Information
- Application Number
- JP2021112077
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-07-06
- Publication Date
- 2025-05-26
- Estimated Expiration
- 2041-07-06
AI Technical Summary
Existing methods for calculating the similarity of media content struggle to effectively capture the changes and transitions in meaning and impression over time, particularly in content with developments and threads that evolve over time.
A content similarity calculation system that extracts semantic waveforms representing changes in semantic items over time, converts these into semantic frequency spectra using Fourier transform, and calculates cosine similarity between these spectra to determine content similarity.
Enables the accurate grasping of changes and transitions in media content meaning and impression over time, allowing for similarity calculations based on development and context, and facilitating effective content search and recommendation systems.
Smart Images

Figure 0007682525000001 
Figure 0007682525000002 
Figure 0007682525000003
Abstract
Description
Technical Field
[0001] The present invention relates to a technique for calculating the similarity of content, and more particularly, to a technique effective for application to a content similarity system and a content search / recommendation system for calculating the similarity of media content having developments and threads that change over time.
Background Art
[0002] In recent years, digital media content (such as novels, music, movies, videos, etc.) has been abundant everywhere on the Internet, and the opportunities for users to enjoy these have been increasing. Here, it is important to establish a method for effectively acquiring and recommending content that matches the user's hobbies, preferences, and intentions. In particular, it is necessary to establish a method for searching and recommending based on the content such as the meaning and impression of media content, rather than simply pattern matching, as one of so-called perceptual search.
[0003] Generally, these media contents have developments and threads that change over time, and their meaning and impression change over time. For example, in novels, even if the expression methods and techniques are different, there are works with similar developments and threads, and in some cases, the author's background can be inferred from the developments and threads. The semantic features of media content are expressed not only by the features of the media content itself but also by the changes in features over time. Therefore, when evaluating a work expressed as media content, it is important to verify not only the overall meaning and impression but also the meaning and impression evoked in each scene.
[0004] In the analysis of general time-series data, it is important to use an appropriate method for conceptually grasping the similarity between two time-series data. This can be roughly divided into a method for directly calculating the similarity between time-series data and a method for converting time-series data into frequency data and calculating the similarity.
[0005] As direct calculation methods for the former, various methods are known, such as DTW (Dynamic Time Warping), ERP (Edit distance with Real Penalty), LCS (Longest Common Subsequence), EDR (Edit Distance on Real sequence), FTSE (Fast Time Series Evaluation), and the like. Also, as methods for converting to frequency data and then calculating for the latter, there are methods using Fourier transform (Non-Patent Document 1, Non-Patent Document 2, etc.) and methods using Haar wavelet.
[0006] In addition, as a technique for considering changes in the semantic content of content in the time series rather than just pattern matching, for example, in Japanese Patent Application Laid-Open No. 2021-9633 (Patent Document 1), the content of the data to be searched is analyzed to calculate an emotion vector in which each emotion item and its emotion amount are associated, and based on the response content of the user to the reference sample, an emotion vector for the event required by the user is calculated. By comparing and calculating the emotion vector calculated for the data to be searched and the emotion vector calculated for the user, an information retrieval method for extracting the data to be searched showing an emotion vector close to the emotion vector calculated for the user is described. It is said that the story property of the time series can be considered by including an item whose emotion amount increases when it is determined that a plurality of specific conditions are satisfied in the time series for the emotion item.
Prior Art Documents
Patent Documents
[0007]
Patent Document 1
Non-Patent Documents
[0008]
Non-Patent Document 1
[0009] For example, by applying conventional calculation methods for similarity of time - series data, including Non - Patent Documents 1 and 2, etc., it is possible to evaluate the similarity of the media content data itself. However, it is difficult to perform a similarity evaluation considering semantic content such as the transition of meaning and impression in the time - series of media content.
[0010] On the other hand, with the technology described in Patent Document 1, it is possible to consider to a certain extent the change in the semantic content of the content in the time - series. However, even when it comes to time - series, it can only consider the order of appearance and frequency of items, and cannot consider the temporal interval during which the semantic content changes.
[0011] Therefore, an object of the present invention is to provide a content similarity calculation system that grasps the changes and transitions of the meaning and impression of media content on the time axis and in time - series, and calculates the similarity based on the development and context of the media content, and a content search and recommendation system using the same.
[0012] The above and other objects and novel features of the present invention will become apparent from the description of this specification and the accompanying drawings.
Means for Solving the Problems
[0013] Among the inventions disclosed in the present application, the outline of typical ones will be briefly described as follows.
[0014] A content similarity calculation system according to a typical embodiment of the present invention is a content similarity calculation system that calculates the similarity between media contents whose contents change over time. For two or more input media contents, a semantic waveform extraction unit that extracts changes in the degree of influence for each of a plurality of semantic items and outputs them as semantic waveform data, and a time-frequency domain conversion unit that obtains and outputs a semantic frequency spectrum by Fourier transform from the semantic waveform data of each of the semantic waveforms of each of the media contents, and for each of the media contents, a semantic frequency spectrum vector having the data of each of the semantic frequency spectra as elements is generated, and a similarity calculation unit that calculates and outputs the cosine similarity between the semantic frequency spectrum vectors.
Effects of the Invention
[0015] Among the inventions disclosed in the present application, the effects obtained by typical ones will be briefly described as follows.
[0016] That is, according to a typical embodiment of the present invention, it is possible to grasp the changes and transitions in the meaning and impression of media contents on the time axis and in time series, and calculate the similarity based on the development and plot of the media contents.
Brief Description of the Drawings
[0017]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Embodiments for Carrying Out the Invention
[0018] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. In all the drawings for explaining the embodiments, the same parts are generally denoted by the same reference numerals, and the repeated explanations thereof are omitted. On the other hand, for the parts explained with reference numerals in a certain drawing, they will not be shown again in the explanations of other drawings, but may be referred to with the same reference numerals.
[0019] (Embodiment 1) <Outline> The content similarity calculation system according to Embodiment 1 of the present invention grasps the changes and transitions (semantic transitions) in time series of the meaning and impression of media content whose content changes in time series, and extracts a meaning waveform expressing this as a waveform. Then, the meaning waveforms extracted for each of a plurality of meaning items in the media content are converted into the frequency domain by Fourier transform to obtain a frequency spectrum, which is expressed as a vector. And for this vector obtained for each media content, by calculating the cosine similarity between the vectors, it is an information processing system for calculating the similarity between media contents based on semantic transitions.
[0020] <System Configuration> FIG. 1 is a diagram showing an overview of a configuration example of a content similarity calculation system 1 according to Embodiment 1 of the present invention. The content similarity calculation system 1 is configured by, for example, a server system such as a server device or a virtual server constructed on a cloud computing service, or a computer system such as a PC (Personal Computer). Then, by a CPU (Central Processing Unit) (not shown), an OS (Operating System), a DBMS (DataBase Management System), middleware such as a Web server program, and software operating thereon, which are expanded from a recording device such as an HDD (Hard Disk Drive) onto a memory, are executed to realize various functions described later related to the similarity calculation of media content.
[0021] The content similarity calculation system 1 has, for example, components such as a content input unit 11, a semantic waveform extraction unit 12, a time-frequency domain conversion unit 13, and a similarity calculation unit 14 implemented as software.
[0022] The content input unit 11 has a function of receiving an input of media content that is a target for calculating similarity. It may directly receive an input of a data file of media content via an input / output device (not shown), or may receive an input by upload via a network 2 such as the Internet. The media content that is the target for similarity calculation does not necessarily have to be of the same type (for example, novels, movies, etc.). In the present embodiment, since the data of media content is not directly compared during similarity calculation, but rather the semantic waveform is extracted and then converted to the frequency domain for similarity calculation, different types of media content can also be subjected to similarity calculation.
[0023] The semantic waveform extraction unit 12 extracts the changes in each time axis and time series for a plurality of feature items from the media content input via the content input unit 11, and based on this, for a plurality of semantic items representing the meaning and impression of the content, it outputs, as data of a semantic waveform, the transition in the time axis and time series of the degree of influence of each. The functions and processing details of the semantic waveform extraction unit 12 will be described later.
[0024] The time-frequency domain conversion unit 13 has a function of obtaining and outputting a semantic frequency spectrum by performing Fourier transform on the data of the semantic waveform related to each semantic item output from the semantic waveform extraction unit 12 to convert the data in the time domain to the data in the frequency domain. The functions and processing details of the time-frequency domain conversion unit 13 will also be described later.
[0025] The similarity calculation unit 14 has a function of calculating and outputting the similarity between each media content for which the similarity is to be calculated, based on the data of the semantic frequency spectrum output from the time-frequency domain conversion unit 13. The calculation of the similarity based on the data of the semantic frequency spectrum is not particularly limited. For example, for each media content to be calculated, a vector (semantic frequency spectrum vector) having the data of the semantic frequency spectrum for each semantic item as elements can be generated, and a method such as calculating the cosine similarity between these vectors can be adopted.
[0026] <Content Similarity Calculation Processing Details> FIG. 2 is a diagram for explaining the outline of the content similarity calculation processing in Embodiment 1 of the present invention. For a plurality of media contents A (41a) and B (41b) to be calculated, which are input via the content input unit 11, semantic waveform extraction processes (S1a, S1b) are respectively performed by the semantic waveform extraction unit 12 to obtain a semantic waveform A (42a) and a semantic waveform B (42b).
[0027] For the semantic waveforms A (42a) and semantic waveforms B (42b), time-frequency domain conversion processing (S2a, S2b) is performed by the time-frequency domain conversion unit 13, respectively, to obtain the semantic frequency spectrum A (43a) and the semantic frequency spectrum B (43b). Then, for the data of the semantic frequency spectrum A (43a) and the semantic frequency spectrum B (43b), the similarity calculation unit 14 vectorizes them respectively to obtain semantic frequency spectrum vectors, and performs similarity calculation processing (S3) to calculate the cosine similarity between these vectors, thereby calculating the similarity 45 between the media content A (41a) and the media content B (41b).
[0028] Hereinafter, the content of each process will be described.
[0029] <Semantic waveform extraction process> FIG. 3 is a diagram for explaining the outline of the semantic waveform extraction process in Embodiment 1 of the present invention. In the present embodiment, a plurality of semantic waveforms 42 are extracted for each semantic item from media content 41 having development or lines, that is, semantic meanings and impressions change over time t. In the example of FIG. 3, it shows that semantic waveforms 42 are extracted for two semantic items of "joy" and "sorrow" respectively. Thus, in the present embodiment, the number of extracted semantic waveforms 42 is the same as the number of semantic items (from the perspective of semantics and impressions) to be evaluated.
[0030] In the semantic waveform extraction process, first, the media content 41 is divided into a plurality of windows 51 separated by a predetermined time width (step 1). Here, the minimum time width capable of expressing semantics is used as the window size, and the media content 41 of interest is divided into a plurality of windows 51 for each window size set from the beginning (or from a specified position).
[0031] Then, a plurality of feature items are respectively extracted from the media content 41 divided into a plurality of windows 51 (step 2). Thereby, it is possible to obtain the change over time of the feature items that appear for each window 51. Here, the feature items can be, for example, sentences, words, images, music, etc. related to meaning and impression that directly appear in the media content 41. Then, each feature item for each window 51 is converted into a semantic item (for example, "joy", "sorrow", etc.) (step 3).
[0032] Then, the magnitude and degree of the influence of each semantic item are quantified for each window 51 (step 4), and a semantic waveform 42 representing the change over time as a waveform is obtained (step 5). In the example of FIG. 3, an example of the semantic waveform 42 obtained for two semantic items, "joy" and "sorrow", in the media content 41 is shown.
[0033] As a method for converting feature items into semantic items in the processes of steps 2 and 3 above, for example, existing general sentiment and emotion analysis methods such as positive / negative judgment and sentiment judgment for sentences and words can be used. Also, for example, it is also possible to use a topic model known as a method of statistical latent semantic analysis of sentences to grasp the change in the share over time of topics (semantic items) in the media content 41.
[0034] Also, for example, it is also possible to use the Media-lexicon Transformation Operator, which is a framework for extracting impression words from media content proposed in the literature (T. Kitagawa and Y. Kiyoki, Y, "Fundamental framework for media data retrieval system using media lexicon transformation operator", Information Modeling and Knowledge Bases, vol. 12, pp. 316-326, 2001). This Media-lexicon Transformation Operator is a mechanism that realizes the extraction of impression words that humans receive from the media content by using research, reviews, statistics, etc. by experts in the field related to the target media content.
[0035] FIG. 4 is a diagram showing an overview of an example of extracting features in time series from media content 41 in Embodiment 1 of the present invention. In this embodiment, the above Media-lexicon Transformation Operator originally targets static one-dimensional data, and this is extended and used on the time axis. For example, a word matrix Y is obtained from the correspondence relationship between the appearance frequencies of words (w 1 , w 2 , …, w m ) related to feature items that appear for each window 51 in the media content 41, and the time (t 1 , t 2 , …) (for example, grasped by the elapsed time from the start of the media content 41) of each window 51. Then, using a conversion matrix T (prepared in advance in dictionary data) consisting of the relationship between the word matrix Y and words (w 1 , w 2 , …, w m ) and semantic items (x 1 , x 2 , …, x n ) indicating meaning and impression, the conversion matrix T and semantic items (x 1 , x 2 , …, xn ) and the time (t 1 , t 2 , …) of each window 51 to convert it into the product with the media content matrix X. From this media content matrix X, the transition of each semantic item over time can be grasped.
[0036] <Time-Frequency Domain Conversion Processing> FIG. 5 is a diagram for explaining the outline of the time-frequency domain conversion processing in Embodiment 1 of the present invention. Here, for the data of the semantic waveforms 42 of each semantic item (in the example of FIG. 5, "joy" and "sorrow"), discrete Fourier transform is performed on each of them by the Fast Fourier Transform (FFT) algorithm to convert from the time domain to the frequency domain, thereby obtaining the semantic frequency spectra for each semantic item (in the example of FIG. 5, the semantic frequency spectrum 43a of "joy" and the semantic frequency spectrum 43b of "sorrow").
[0037] <Similarity Calculation Processing> In the similarity calculation processing, for each media content 41, the semantic frequency spectra for each semantic item obtained by the above series of processes are vectorized with the elements of each column to obtain a semantic frequency spectrum vector 44. Then, the similarity is obtained by calculating the cosine similarity of the semantic frequency spectrum vectors 44 for each media content 41. Regarding the similarity, the value of the cosine similarity may be used as it is, or by comparing this value with a predetermined reference value, it may be used to determine whether they are similar or not, or to evaluate the degree of stepwise similarity.
[0038] <Conclusion> As described above, according to the content similarity calculation system 1 which is Embodiment 1 of the present invention, it is possible to grasp the transition of meaning and impression in the media content 41 on the time axis and in time series, and calculate the similarity between contents based on the development and context of the media content 41.
[0039] Also, instead of directly comparing the data of media content 41 during similarity calculation, a semantic waveform 42 is extracted from the media content 41, and similarity calculation is performed based on the semantic frequency spectrum 43 obtained by converting this into the frequency domain. Therefore, it is possible to perform similarity calculation even for different types of media content 41, such as novels and movies.
[0040] Furthermore, during similarity calculation, for each media content 41, a simple method is adopted where a semantic frequency spectrum vector 44 is generated by vectorizing a plurality of semantic frequency spectra 43, and the cosine similarity between these vectors is calculated to obtain the similarity. By doing so, it is possible to calculate the similarity between media contents 41 at low cost.
[0041] (Embodiment 2) The content search and recommendation system according to Embodiment 2 of the present invention is an information processing system that performs search and recommendation of media content using the content similarity calculation system of Embodiment 1 described above.
[0042] <System Configuration> FIG. 6 is a diagram showing an overview of a configuration example of the content search and recommendation system according to Embodiment 2 of the present invention. The content search and recommendation system 3 is configured by, for example, a server system such as a server device or a virtual server constructed on a cloud computing service, or a computer system such as a PC. Then, by a CPU (not shown) executing middleware such as an OS, a DBMS, and a Web server program developed from a recording device such as an HDD onto a memory, and software operating thereon, various functions described later related to search and recommendation of media content based on conditions specified by a user are realized.
[0043] The content search and recommendation system 3 includes components such as a conditional content input unit 31, a content extraction unit 32, a content similarity calculation unit 33, and a result content output unit 34, which are implemented as software, for example. It also has a content database (DB) 35 that stores multiple media contents. The content search and recommendation system 3 extracts media contents similar to the media content specified by the user as search conditions from the media contents stored in the content DB 35 and outputs them as search results or recommended contents.
[0044] The conditional content input unit 31 has a function of receiving input of a comparison target media content that serves as a condition when performing media content search or recommendation. It may directly receive input of a data file of media content via an input / output device (not shown), or may receive input by upload via a network 2 such as the Internet.
[0045] The content extraction unit 32 has a function of sequentially extracting the media contents stored in the content DB 35 as search targets. It may extract all media contents sequentially, or may narrow down the range of media contents to be extracted from the content DB 35 based on the content, metadata, and other information of the media content with the search conditions input via the conditional content input unit 31 and then extract them sequentially.
[0046] The content similarity calculation unit 33 has a function of calculating the similarity between media contents, taking as input the media content with the search conditions input via the conditional content input unit 31 and the media contents sequentially extracted from the content DB 35 by the content extraction unit 32. In implementation, for example, the content similarity calculation system 1 (more precisely, each component implemented as software) described in the above-described Embodiment 1 can be used, and it is assumed that such a configuration is also adopted in this embodiment.
[0047] Based on the similarity information calculated by the content similarity calculation unit 33, the result content output unit 34 outputs to the user a certain number (or those with a similarity equal to or higher than a certain value) of media contents in the content DB 35 in descending order of the similarity to the media content of the search condition as search results or recommended contents. The output method is not particularly limited, and it may be displayed on a display (not shown) provided in the content search and recommendation system 3, or it may be taken out as a data file or downloaded by the user via the network 2.
[0048] <Conclusion> As described above, according to the content search and recommendation system 3 which is the second embodiment of the present invention, it is possible to extract media contents similar to the media content specified as the search condition by the user from the media contents stored in the content DB 35 and output them as search results or recommended contents.
[0049] As described above, the invention made by the present inventor has been specifically described based on the embodiments. However, it goes without saying that the present invention is not limited to the above embodiments and can be variously modified without departing from the gist thereof. Also, the above embodiments have been described in detail for easy understanding of the present invention, and are not necessarily limited to those having all the configurations described. Also, a part of the configuration of one embodiment can be replaced with the configuration of another embodiment, and the configuration of another embodiment can be added to the configuration of one embodiment. Also, it is possible to add, delete, or replace other configurations for a part of the configuration of each embodiment.
[0050] In addition, each of the above-described configurations, functions, processing units, processing means, etc. may be realized in hardware by designing a part or all of them, for example, by means of an integrated circuit. Further, each of the above-described configurations, functions, etc. may be realized in software by a processor interpreting and executing a program for realizing each function. Information such as a program, table, file, etc. for realizing each function can be placed in a recording device such as a memory, hard disk, SSD (Solid State Drive), or a recording medium such as an IC card, SD card, DVD.
[0051] In addition, in each of the above figures, control lines and information lines show those considered necessary for explanation, and do not necessarily show all control lines and information lines in implementation. Practically, it may be considered that almost all configurations are interconnected.
Industrial Applicability
[0052] The present invention can be used in a content similarity system for calculating the similarity of media content having developments and streaks that change over time, and a content search / recommendation system.
Explanation of Signs
[0053] 1... Content similarity calculation system, 2... Network, 3... Content search / recommendation system, 11... Content input unit, 12... Meaning waveform extraction unit, 13... Time / frequency domain conversion unit, 14... Similarity calculation unit, 31... Condition content input unit, 32... Content extraction unit, 33... Content similarity calculation unit, 34... Result content output unit 34, 35... Content DB, 41... Media content, 41a... Media content A, 41b... Media content B, 42... Meaning waveform, 42a... Meaning waveform A, 42b... Meaning waveform B, 43... Meaning frequency spectrum, 43a... Meaning frequency spectrum A, 43b... Meaning frequency spectrum B, 44... Meaning frequency spectrum vector, 45... Similarity 51... Window
Claims
1. A content similarity calculation system for calculating the similarity between media contents whose content changes over time, comprising: a semantic waveform extraction unit that extracts changes in the degree of influence for each of a plurality of semantic items for each of the input two or more media contents and outputs them as semantic waveform data; a time-frequency domain conversion unit that obtains and outputs a semantic frequency spectrum by Fourier transform from the semantic waveform data of each of the semantic waveforms of each of the media contents; a similarity calculation unit that generates a semantic frequency spectrum vector having the data of each of the semantic frequency spectra as elements for each of the media contents, and calculates and outputs the cosine similarity between the semantic frequency spectrum vectors; A content similarity calculation system having the above.
2. In the content similarity calculation system according to Claim 1, the semantic waveform extraction unit Based on a word matrix consisting of the correspondence between the number of occurrences of words related to characteristic items in each window obtained by dividing the media content by a predetermined time width and the time of each window, a word related to the characteristic item, and a conversion matrix consisting of the correspondence between each semantic item, a media content matrix is obtained such that the product with the conversion matrix becomes the word matrix, and the data of each semantic waveform is obtained based on the data of the media content matrix. A content similarity calculation system.
3. A content search and recommendation system for searching for media contents similar to a specified second media content from among a content recording unit in which a plurality of first media contents whose content changes over time are recorded, comprising: a content similarity calculation unit comprising the content similarity calculation system according to Claim 1; a result content output unit that outputs, as a search result, the first media content whose similarity satisfies a predetermined condition based on the similarity between each of the first media contents and the second media content calculated by the content similarity calculation unit; A content search and recommendation system having the above.
Citation Information
Patent Citations
Method and device for calculating similarity, program and recording medium
JP2004046370A
User support method, device, and program
JP2008269065A
Information retrieval method and information retrieval device
JP2021009633A