Video processing device and video processing program
The video processing device and program effectively detect scene changes by analyzing video data features and elements, ensuring accurate scene change detection and summary generation.
Patent Information
- Application Number
- JP2024028135
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-28
- Publication Date
- 2025-09-09
AI Technical Summary
Existing video processing techniques struggle to accurately detect scene changes.
A video processing device and program that analyze video data at predetermined intervals, calculate image features, detect and extract elements, and determine scene changes based on feature and element differences, using methods such as color histograms, shape features, HoG features, and CNN features, to identify scene change points.
Accurately detects scene changes by identifying feature and element variations, enabling the generation of summaries between scene change points through text conversion and attribute assignment.
Smart Images

Figure 2025130819000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a technique for processing video images. [Background technology]
[0002] A known prior art technique for processing video is a personal video recorder that includes a processing unit, a first communication interface, a second communication interface, a data storage device, recorded program data stored in the data storage device, the recorded program data defining the recording of a television program, event data stored in the data storage device, the event data defining a plurality of events and a corresponding time for each event, index data stored in the data storage device, the corresponding time for each event and a key frame byte offset for each corresponding time, and playback logic accessible by the processing unit, the playback logic (i) determining whether an event has occurred, (ii) dividing the recorded program data into scenes, and (iii) skipping at least a portion of at least one scene in response to a command from a user (see Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Special Publication No. 2007-502080 Summary of the Invention [Problem to be solved by the invention]
[0004] However, the above-described conventional techniques have a problem in that scene changes may not be detected accurately.
[0005] Therefore, in one aspect, the present invention aims to provide a technique that can accurately detect scene changes. [Means for solving the problem]
[0006] In order to solve the above problem, a video processing device according to the present invention includes means for acquiring video data of a plurality of moments at predetermined time intervals for video data to be processed, means for calculating, for each of the video data of the plurality of moments, a feature amount of an image generated by the video data of that moment, means for detecting and extracting elements included in an image generated by the video data of that moment for each of the video data of the plurality of moments, and means for determining, when the feature amounts of the images of the video data of the moments preceding and succeeding in time series are different, the interval between the video data of the moments preceding and succeeding in time series as a candidate for a scene change point and selecting one of the video data of the moments preceding and succeeding in time series. The video data system may have a means for assigning a feature change flag, a means for determining, when the elements contained in the images of the video data of successive moments in time series are different, that the image between the video data of the successive moments is a candidate for a scene change point and assigning an element change flag to one of the video data of the successive moments, and a means for determining that the video data of the instant to which at least one of the feature change flag and the element change flag is assigned is a scene change point, and the elements contained in the images may be at least one of people, captions, caption genres, music, audio genres, and locations.
[0007] In the video processing device according to the present invention, the feature of the image may be any one of a color histogram, a shape feature, an edge feature, an HoG feature, and a CNN feature.
[0008] The video processing device according to the present invention may be configured such that the elements included in the image include at least one of the genre of the telop and the genre of the audio as essential elements.
[0009] The video processing device of the present invention may further have means for generating a summary for each section in time series from a certain scene change point to the next scene change point among the scene change points determined to be the scene change points, and the summary may be the result of at least one of processes including converting audio to text, converting subtitles to text, converting objects to text, and assigning attributes based on text data generated by converting the audio to text, the subtitles to text, and the objects to text.
[0010] Furthermore, a video processing program according to the present invention includes a process of acquiring video data of a plurality of moments at predetermined time intervals for video data to be processed, a process of calculating, for each of the video data of the plurality of moments, a feature amount of an image generated by the video data of that moment, a process of detecting and extracting, for each of the video data of the plurality of moments, an element included in an image generated by the video data of that moment, and a process of determining, when the feature amounts of the images of the video data of the moments preceding and succeeding in time series are different, the video data of the moments preceding and succeeding as a candidate for a scene change point and setting a feature amount change flag for one of the video data of the moments preceding and succeeding a process of determining, if the elements contained in the image of each of the video data of the preceding and succeeding moments in time series are different, the interval between the video data of the preceding and succeeding moments as a candidate for a scene change point and assigning an element change flag to one of the video data of the preceding and succeeding moments; and a process of determining that the video data of the moment to which at least one of the feature change flag and the element change flag is assigned is a scene change point, and the elements contained in the image may be at least one of a person, a caption, a caption genre, music, an audio genre, and a location.
[0011] In the video processing program according to the present invention, the image feature may be any one of a color histogram, a shape feature, an edge feature, an HoG feature, and a CNN feature.
[0012] The video processing program according to the present invention may be configured such that the elements included in the image include at least one of the genre of the telop and the genre of the audio as essential elements.
[0013] The video processing program of the present invention may further cause the information processing device to execute a process of generating a summary for each section in a time series from a certain scene change point to the next scene change point among the scene change points determined to be the scene change points, and the summary may be the result of at least one process of converting audio into text, converting subtitles into text, converting objects into text, and assigning attributes based on the text data generated by converting the audio into text, the subtitles into text, and the objects into text. [Effects of the Invention]
[0014] According to one aspect, the present invention makes it possible to accurately detect scene changes. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a schematic diagram illustrating the overall configuration of a video processing system including a video processing device according to an embodiment of the present invention. [Figure 2] 2 is a schematic diagram illustrating the hardware configuration and functional configuration of a video processing server of the video processing system of FIG. 1. [Figure 3] 3 is a flowchart showing a processing procedure in the video processing server of FIG. 2. [Figure 4] FIG. 10 is a diagram showing an example of genres assigned to text data. DETAILED DESCRIPTION OF THE INVENTION
[0016] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. In the following embodiments, a video processing device according to the present invention is incorporated into a video processing system as a video processing server.
[0017] (Video Processing System) FIG. 1 is a schematic diagram of the overall configuration of a video processing system 100 including a video processing server 1 according to an embodiment, as an example of a specific configuration of a video processing device according to the present invention.
[0018] The video processing server 1 is in the form of a server and is configured, for example, by a server computer. The video processing server 1 may be realized by multiple server computers. For example, the video processing server 1 may be realized by multiple server computers located in different locations working together. The video processing server 1 may include a web server.
[0019] The video processing server 1 can communicate with other servers (for example, a cloud server 2 in the illustrated example) via a communication network 9. That is, the video processing server 1 is connected so as to be able to communicate with external devices via the communication network 9.
[0020] The communication network 9 is a communication network consisting of, for example, the Internet, various wireless communication lines, wired communication lines, etc., and may include a LAN (Local Area Network), a WAN (Wide Area Network), an intranet, Ethernet (registered trademark), etc. The communication network 9 may include a wireless network or a wired network. At least a cloud server 2 is connected to the communication network 9. Part or all of the communication network 9 may be realized by a P2P network.
[0021] The cloud server 2 is in the form of a server and is configured, for example, by a server computer. The cloud server 2 may be realized by multiple server computers. For example, the cloud server 2 may be realized by multiple server computers located in different locations working together. The cloud server 2 may include a web server.
[0022] The cloud server 2 is capable of communicating with an external device (for example, the video processing server 1 in the illustrated example) via the communication network 9. That is, the cloud server 2 is connected so as to be able to communicate with the external device via the communication network 9.
[0023] The cloud server 2 stores data and information required when the video processing server 1 performs processing related to scene change detection. In this embodiment, the cloud server 2 stores a video database 21.
[0024] The video database 21 stores and records video data that is the subject of processing by the video processing server 1 for scene change detection.
[0025] The video data in this invention is a concept that includes various types of video data, and the video format and the manner of providing to viewers are arbitrary. The video data is typically data that is transmitted to an unspecified number of people, and includes, for example, terrestrial broadcast video data, satellite broadcast video data, and internet distribution video data.
[0026] In the video database 21, video data is accumulated and recorded in a manner that allows each group of video data, such as each program or each video content, to be distinguished from each other when the video data is broadcast or distributed.
[0027] The video data may include audio, music, and captions associated with the video. Examples of audio associated with the video include narration, dialogue, conversation, reading of a script, commentary, etc. Examples of music associated with the video include background music (music played in the background of the video and audio) and music performed to accompany the video of a performance.
[0028] (Video processing server) FIG. 2 is a schematic diagram of the hardware configuration and functional configuration of the video processing server 1. As shown in FIG.
[0029] In the example shown in the figure, the video processing server 1 includes a control unit 11, a storage unit 12, an input unit 13, a communication unit 14, and a drive device 15, and these components are electrically connected via a bus 19 inside the video processing server 1. The video processing server 1 may further include components other than the above-mentioned components. The video processing server 1 may further include, for example, a display unit (not shown).
[0030] The control unit 11 controls various types of arithmetic processing (especially processing related to scene change detection) and overall operations related to the video processing server 1. The control unit 11 may be configured to include, for example, a central processing unit (CPU).
[0031] The control unit 11 realizes various functions related to the video processing server 1 by reading out predetermined programs (including video processing programs) stored in the storage unit 12 and executing processes in accordance with the programs. In other words, the various functions related to the video processing server 1 can be executed as each functional unit included in the control unit 11 by specifically implementing information processing by the software (including video processing programs) stored in the storage unit 12 by the control unit 11, which is an example of hardware. The control unit 11 is not limited to being a single unit, and may be configured to have multiple control units 11 for each function. Furthermore, the control unit 11 may be a combination of multiple control units.
[0032] The storage unit 12 may be configured with any storage medium and stores various programs (including video processing programs), information, and data. The storage unit 12 may be configured with, for example, a storage device such as an HDD (Hard Disk Drive) or a Solid State Drive (SSD) that stores various programs and the like related to the video processing server 1 executed by the control unit 11, a memory such as a read-only memory (ROM) that similarly stores various programs and the like, or a random access memory (RAM) that stores temporarily required information, data, and the like related to program calculations, or a combination thereof.
[0033] The input unit 13 functions as an input interface that accepts input operations by an operator. The input unit 13 may be included in the housing of the video processing server 1 (in other words, integrated with the housing), or may be attached externally. The input unit 13 may be configured, for example, by various input keys, operation keys, or a mouse for inputting numerical values and data and for issuing instructions for operations and actions.
[0034] The communication unit 14 has a function as a communication interface for the communication network 9, connecting the video processing server 1 and the communication network 9 so that they can communicate with each other. The communication unit 14 performs communication processing with an external device connected via the communication network 9 so that they can communicate with each other.
[0035] The drive device 15 reads out a program (including a video processing program) from a recording medium, such as a flexible disk, and stores it in the storage unit 12.
[0036] The recording medium stores predetermined programs (including video processing programs). The programs stored on this recording medium are installed in the video processing server 1 via the drive device 15. The installed predetermined programs can be executed by the video processing server 1.
[0037] In the example shown in the figure, the various processes described below can be realized by having the video processing server 1 execute a program (including a video processing program). Alternatively, the program can be recorded on a recording medium, and the video processing server 1 can read the program from the recording medium to realize the various processes described below. Various types of recording media can be used. For example, the recording media may be recording media that record information optically, electrically, or magnetically, such as a CD (Compact Disc)-ROM, a flexible disk, or a magneto-optical disk, or may be semiconductor memory that records information electrically, such as a ROM or flash memory.
[0038] (Processing content) FIG. 3 is a flowchart showing a processing procedure in the video processing server 1 according to the embodiment.
[0039] The video processing server 1 according to the embodiment includes a still image generation unit 113 as means for acquiring video data of a plurality of moments at a predetermined time interval for video data to be processed, a feature amount calculation unit 114 as means for calculating, for each of the video data of the plurality of moments, a feature amount of an image generated by the video data of that moment, an element extraction unit 115 as means for detecting and extracting elements included in the image generated by the video data of that moment for each of the video data of the plurality of moments, and a feature amount change flag as means for determining, when the feature amounts of the images of the video data of successive moments in time series are different, the video data of the successive moments as candidates for a scene change point and assigning a feature amount change flag to the video data of the later moment out of the video data of the successive moments. the feature amount determination unit 116; an element determination unit 117 as means for determining, when the elements contained in each image of video data for successive moments in time series are different, that the data between these moments is a candidate for a scene change point and assigning an element change flag to the video data for the later moment of the video data for the successive moments; a change point identification unit 118 as means for determining that video data for a moment to which at least one of a feature amount change flag and an element change flag is assigned is a scene change point; and a summary generation unit 119 as means for generating a summary for each section from a scene change point to the next scene change point in time series among the scene change points determined to be scene change points.
[0040] Furthermore, the video processing program according to the embodiment includes a process of acquiring video data of a plurality of moments at predetermined time intervals for video data to be processed (steps S2 and S3), a process of calculating, for each of the video data of the plurality of moments, a feature amount of an image generated by the video data of that moment (step S4), a process of detecting and extracting, for each of the video data of the plurality of moments, an element included in the image generated by the video data of that moment (step S5), and a process of determining, when the feature amounts of the images of the video data of the preceding and succeeding moments in time series are different, as a candidate for a scene change point between the video data of the preceding and succeeding moments and assigning a feature amount change flag to the video data of the succeeding moment among the video data of the preceding and succeeding moments (step S6). The video processing server 1 as an information processing device is caused to perform the following steps: if the elements contained in each image of the video data for successive moments in the time series are different, determine the interval between these video data for successive moments as a candidate for a scene change point and assign an element change flag to the video data for the later moment of the video data for successive moments (step S7); determine that the video data for the moment to which at least one of the feature change flag and the element change flag is assigned is a scene change point (step S8); and generate a summary for each section from a certain scene change point to the next scene change point in the time series of the scene change points determined to be scene change points (step S9).
[0041] The control unit 11 of the video processing server 1 includes an instruction receiving unit 111, a data acquisition unit 112, a still image generation unit 113, a feature calculation unit 114, an element extraction unit 115, a feature determination unit 116, an element determination unit 117, a switching point identification unit 118, and a summary generation unit 119.
[0042] Each functional unit of the control unit 11 is generated by the control unit 11 reading and executing a predetermined program (including a video processing program) stored in the storage unit 12. Note that "transmission" in the description of the present invention includes the exchange of information and data that is performed by storing in the storage unit 12 and reading from the storage unit 12.
[0043] When the instruction receiving unit 111 receives an instruction to start video processing, it notifies the data acquiring unit 112 of the start of processing (step S1).
[0044] The instruction to start video processing may be input via the input unit 13 of the video processing server 1 and transmitted to the instruction receiving unit 111 via the bus 19, or may be input from an external device via the communication unit 14 and transmitted to the instruction receiving unit 111 via the bus 19. The instruction to start video processing may be automatically generated within the video processing server 1 periodically or in response to a predetermined trigger.
[0045] When an instruction to start video processing is transmitted from instruction receiving unit 111, data acquiring unit 112 acquires video data from video database 21 stored in cloud server 2 via communication unit 14 (step S2).
[0046] The processing by the video processing server 1 is performed for each piece of video data (for example, one program or one piece of video content) accumulated and recorded in the video database 21. In the description of the present invention, a certain piece of video data acquired from the video database 21 by the data acquisition unit 112 and subjected to video processing is referred to as "video data to be processed." In the following description, processing on a certain piece of video data as the video data to be processed is described.
[0047] The data acquisition unit 112 transmits the video data acquired from the video database 21 to the still image generation unit 113 as video data to be processed.
[0048] The still image generating unit 113 acquires a plurality of instantaneous video data at predetermined time intervals from the video data to be processed transmitted from the data acquiring unit 112 (step S3).
[0049] Specifically, the still image generation unit 113 targets video data (moving image) composed of a series of multiple images (still images) at a predetermined frame rate (fps), extracts images (still images) from the multiple images (still images) at predetermined time intervals (referred to as "still image generation intervals"), and stores them in the memory unit 12 as still image data.
[0050] The still image generation interval is not limited to a specific time length, but is set to an appropriate time length, taking into consideration, for example, the genre of the program of the video data to be processed, etc. Specifically, the still image generation interval may be set to a time length of about 0.5 to 5 seconds, for example.
[0051] The still image generating unit 113 stores a set of still image data acquired (extracted) from the video data to be processed in the storage unit 12 as a still image database 121.
[0052] Here, a unique identifier for distinguishing and identifying one piece of video data may be assigned to each piece of video data, and a unique identifier for distinguishing and identifying one piece of still image data obtained (extracted) from the video data may be assigned to each piece of still image data. Then, a certain piece of still image data in the video data to be processed may be identified by a combination of the identifier assigned to each piece of video data and the identifier assigned to each piece of still image data (referred to as a "still image data ID").
[0053] It is preferable that the still image data IDs are assigned to each piece of still image data in consecutive numbers in the chronological order of the video data to be processed from which the still image data was acquired. In this embodiment, the still image data IDs are assigned to each piece of still image data in consecutive numbers in the chronological order of the video data to be processed from which the still image data was acquired, so that the chronological order of each piece of still image data in the video data to be processed can be determined and identified by referring to the still image data IDs.
[0054] Each still image data ID may be associated with the elapsed time from the start of the original video data to be processed. In this case, the chronological order of each still image data piece in the video data to be processed can be determined and identified by referring to the still image data ID assigned to each still image data piece and the elapsed time associated with each still image data ID.
[0055] Combination data (plural sets) of still image data IDs and still image data is accumulated and recorded in the still image database 121. Alternatively, combination data (plural sets) of still image data IDs, elapsed time, and still image data may be accumulated and recorded in the still image database 121.
[0056] The feature amount calculation unit 114 calculates the feature amount of an image generated from each of the still image data stored in the still image database 121 (step S4). An image generated from still image data is called a "still image."
[0057] The feature amount of a still image is a numerical value that quantitatively represents the features and characteristics related to the color, shape, brightness, texture, etc. contained in the still image.
[0058] It is preferable that the feature calculation unit 114 calculates an appropriate index as a feature of a still image to compare images and evaluate the degree of similarity or match. Specifically, for example, it may calculate at least one of the following items A to E.
[0059] a) Color Histograms A color histogram is an index that represents the color distribution within an image, and is an index that quantifies the color distribution by generating a histogram in, for example, an RGB color space or an HSV color space.
[0060] a) Shape Features Shape features are indices that represent the shapes of objects and contours in an image, and describe edges, area boundaries, object contours, etc., and are indices that quantify, for example, the perimeter, area, aspect ratio, etc. of contours.
[0061] C) Edge Features The edge feature amount is an index used to detect edges (i.e., boundaries of objects) in an image, and is an index that quantifies, for example, edge strength, edge direction, edge distribution, and the like.
[0062] d) HoG features (Histograms of oriented Gradients) HoG features are indices that capture the shape of an object and are indices that represent the brightness gradient as a histogram for each angle.
[0063] E) CNN Features (Convolutional Neural Network Features) A convolutional neural network (CNN) may be used to extract features from images and learn the feature quantities.
[0064] The feature calculation unit 114 may calculate audio feature amounts instead of or in addition to the still image feature amounts. Specifically, for each of a plurality of still image data stored in the still image database 121, the feature calculation unit 114 may calculate audio feature amounts for each sentence (or for a predetermined time length, for example, about 10 seconds) including the audio uttered at the moment of the still image data in the original video data to be processed.
[0065] Speech features (also called "acoustic features") are numerical values that quantitatively represent features and characteristics related to amplitude, frequency, temporal characteristics, energy, spectral characteristics, formants, etc. extracted from a speech signal.
[0066] It is preferable that the feature calculation unit 114 calculates an appropriate index as a feature of the voice to compare voices and evaluate the degree of similarity or agreement, and specifically, for example, it may calculate at least one of the following (a) and (b).
[0067] F) Mel spectrum The Mel spectrum is an index that converts an audio signal into different frequency components and quantifies information such as pitch and timbre.
[0068] MFCC (Mel Frequency Cepstrum Coefficients) MFCC (Mel Frequency Cepstrum Coefficients) are coefficients obtained by applying a discrete cosine transform to a logarithmic Mel spectrogram to convert frequency components into principal components.
[0069] The feature calculation unit 114 associates at least one of the still image features and audio features calculated for each still image data stored in the still image database 121 (referred to as a "determination feature") with the still image data ID of each still image data, and stores this in the storage unit 12 as a feature database 122. In other words, the feature database 122 stores and records combination data (including multiple pairs) of still image data IDs and determination features.
[0070] The element extraction unit 115 detects and extracts elements included in an image (that is, a still image) generated from each of a plurality of still image data stored in the still image database 121 (step S5).
[0071] Specifically, the element extraction unit 115 detects and extracts at least one element from among the following 1 to 5 elements.
[0072] 1) Person The element extraction unit 115 uses face recognition technology and person determination technology to detect and extract people included in still images. At this time, each person included in a still image of the video data to be processed is recognized and determined separately, and a unique identifier is assigned to each person to distinguish and identify each person. If the still image does not include a person, an identifier (or code, flag, or label) equivalent to "no person" is assigned.
[0073] The face recognition technology and person determination technology used in the element extraction unit 115 are not limited to a specific method, but may use, for example, OpenAI's large-scale language model "GPT-4" and generation AI "ChatGPT", Amazon Rekognition's Recognizing Celebrities function, Google Cloud Vision API's Celebrity recognition function, or Azure AI Vision's Face function.
[0074] 2) Subtitles The element extraction unit 115 uses character recognition technology to detect and extract the captions contained in the still image. Specifically, it converts the captions contained in the still image into text. If the still image does not contain captions, an identifier (or code, flag, or label) corresponding to "no captions" is assigned to the image.
[0075] Depending on the number of captions included in one still image, text data corresponding to one sentence or character string, or text data corresponding to multiple sentences or character strings, is assigned to a certain still image.
[0076] The character recognition technology used in the element extraction unit 115 is not limited to a specific method, but may be, for example, OpenAI's large-scale language model "GPT-4" and generation AI "ChatGPT", Amazon Textract, the OCR function of Google Cloud Vision API, or the OCR function of Microsoft Azure AI Vision.
[0077] 3) Music The element extraction unit 115 uses a music determination technique to determine the name of the music playing at the moment of the still image in the original video data to be processed. The determined music title may be converted into a unique identifier for identifying and distinguishing these music titles from one another and assigned to the still image. If no music is playing at the moment of the still image, an identifier (or code, flag, or label) corresponding to "no music" is assigned to the still image.
[0078] The music determination technology used in the element extraction unit 115 is not limited to a specific method, but may be, for example, the music recognition technology of Gracenote or the music recognition technology of ACRCloud.
[0079] 4) Audio The element extraction unit 115 uses speech recognition technology (Speech to Text) to detect and determine the sound uttered at the moment of the still image in the original video data to be processed. Specifically, the sound is converted into text. For example, text data of a sentence unit including words uttered at the moment of the still image is associated with the still image as the sound uttered at the moment of the still image in the original video data to be processed.
[0080] The speech recognition technology used in the element extraction unit 115 is not limited to a specific method, but may be, for example, Amazon Transcribe, AmiVoice Cloud Platform, OpenAI's speech recognition model "Whisper", Google Cloud's Speech-to-Text function, or Microsoft Azure's speech-to-text conversion function.
[0081] The element extraction unit 115 further performs morphological analysis, vector representation, machine learning, and deep learning on the text data of the speech in sentence units associated with the still image, and assigns a genre to the text data of the speech content.
[0082] The element extraction unit 115 may also perform machine learning and deep learning on the text data (2 above), which is character information extracted from the captions contained in the still image, and assign a genre to the text data of the content displayed on the screen.
[0083] That is, the element extraction unit 115 may perform machine learning or deep learning on at least one of the text data of the spoken content and the text data of the displayed content (specifically, the caption) and assign a genre to it.
[0084] With regard to the element extraction unit 115, machine learning for assigning a genre to at least one of the text data of the utterance content and the text data of the display content may be performed, for example, by the following procedure.
[0085] As a preliminary preparation, machine learning is performed using a dataset consisting of a set of multiple combinations of nouns and the genres of the nouns as training data, and a trained model is created that takes nouns (words) as input and outputs the genres of the nouns (words). The created trained model is incorporated into the element extraction unit 115.
[0086] As merely an example, a dataset used as training data when creating a trained model may include a combination of a politician's name and a political genre classification, a famous business leader's name and a business genre classification, or an athlete's name and a sports genre classification.
[0087] The text data of the spoken content and the text data of the displayed content are broken down into parts of speech such as nouns, verbs, and adjectives on a word-by-word basis using natural language processing (NLP). Alternatively, nouns (on a word-by-word basis) are extracted from the text data of the spoken content and the text data of the displayed content using natural language processing (NLP).
[0088] A trained model is used to assign a genre to each of the decomposed words that are determined to be nouns, or to each extracted noun.
[0089] Then, the number of times each genre appears is counted for each piece of text data corresponding to one or more sentences or character strings associated with the still image, and the genre with the most appearances is assigned as the genre of the still image.
[0090] The genres assigned to the text data of the spoken content and the text data of the displayed content are not limited to a specific type, but may be assigned in advance as shown in FIG. 4, for example.
[0091] If no sound is being emitted at the moment of the still image, an identifier (or code, flag, or label) corresponding to "no sound" is assigned.
[0092] 5) Location The element extraction unit 115 uses object recognition techniques to determine the locations that are represented in the still image (in other words, as the background of the still image).
[0093] The object recognition technology used in the element extraction unit 115 is not limited to a specific method, but may be, for example, Amazon Rekognition, OpenAI's large-scale language model "GPT-4" and generative AI "ChatGPT", Google Cloud Vision API, or the object detection function of Microsoft Azure AI Vision.
[0094] Locations depicted in still images include, for example, indoor locations such as studios, theaters, and halls, and outdoor locations such as street corners, outdoor facilities, and natural environments (the sea, mountains).
[0095] The element extraction unit 115 recognizes and judges each location shown in a still image of the video data to be processed separately, and assigns a unique identifier to each location for distinguishing and identifying the locations. For example, when indoor sets (e.g., studios) are different, each set is recognized and judged separately, and a unique identifier is assigned to each set for distinguishing and identifying the sets, and when outdoor scenes (e.g., street corners) are different, each scene is recognized and judged separately, and a unique identifier is assigned to each scene for distinguishing and identifying the scenes.
[0096] The target of object recognition by the element extraction unit 115 is not limited to only places, but may include at least one of places, objects, activity details, animals, products, vehicles, buildings, signs, logos, etc. The target recognized as an object by the element extraction unit 115 (in other words, the result of object recognition by the element extraction unit 115) is referred to as an "image element."
[0097] The element extraction unit 115 performs processing on at least one of elements 1 to 5 above for each still image data stored in the still image database 121, and then, depending on the processing performed, associates the extracted person identifier, the textified caption, the textified and determined caption genre, the determined music title (or the identifier of the music title), the determined audio genre, and the determined location identifier with the still image data ID of each still image data, and stores them in the memory unit 12 as element database 123.
[0098] Therefore, the element database 123 stores and records combination data (including multiple sets) of still image data ID and at least one of a person identifier, subtitles (text), subtitle (text) genre, music title, audio genre, and location identifier.
[0099] At least one of the elements 1 to 5 above that is processed by the element extraction unit 115 is referred to as an “image element.” In other words, the element database 123 stores and records combination data (including multiple pairs) of still image data IDs and image elements.
[0100] At least one of the elements 1 to 5 above may be set as a required image element. In this invention, it is preferable that the image elements include the above-mentioned telop genre and audio genre, and at least one of the telop genre and audio genre may be set as a required image element.
[0101] The feature determination unit 116 reads out the combination data (note that there are multiple sets) of still image data IDs and determination features recorded in the feature database 122 for the video data to be processed, and compares the determination features of each still image data that comes before and after in chronological order (note that the order in the chronological order is specified, for example, by the still image data ID) with each other (step S6).
[0102] When the judgment features of successive still image data in the time series are different, the feature determination unit 116 determines that the space between these successive still image data is a candidate for a scene change point, and assigns a feature change flag to the later still image data in the time series (or the earlier still image data in the time series) of these successive still image data (specifically, for example, the value of the feature change flag, which has an initial value of 0, is changed to 1).
[0103] A threshold value (referred to as a "feature threshold") relating to the degree of difference in the determination feature for determining whether or not a scene is a candidate for a scene change point may be set in advance. The feature threshold value differs depending on the type of determination feature (e.g., A to G above), and is not limited to a specific value, but is set to an appropriate value in consideration of whether or not a scene change point can be accurately estimated depending on the genre of the program.
[0104] For example, since a large caption (specifically, a corner title or a topic title) is often displayed on the screen (specifically, in the center or bottom of the screen) at a scene change point, it may be determined that a scene has changed if the size of the caption displayed on the screen (specifically, in the center or bottom of the screen) exceeds a certain value. For this reason, the feature threshold may be set to a value corresponding to a state in which a caption that exceeds a predetermined size (in other words, is conspicuous) is displayed on the screen (specifically, in the center or bottom of the screen).
[0105] Furthermore, since the same side superimposition (in other words, side caption) is often displayed continuously in the upper right or upper left of the screen during the same segment or the same topic, if a caption that has been displayed in a predetermined position for a predetermined length of time is no longer displayed, it may be determined that a scene has changed. For this reason, the feature threshold may be set to a value (or content) that corresponds to the state in which a caption that has been displayed in a predetermined position for a predetermined length of time is no longer displayed.
[0106] In addition, since a scene often changes when a speaker changes, if the speaker changes and the voice characteristics change, it may be determined that a scene has changed. For this reason, the feature threshold may be set to a value (or content) corresponding to the state in which the voice features have changed.
[0107] Furthermore, since the same nouns or nouns with similar meanings (especially proper nouns) are repeatedly mentioned within the same scene, it can be determined that the scene is the same as long as the utterance content contains the same nouns or nouns with similar meanings (especially proper nouns).For this reason, the feature threshold may be set to a value (or content) corresponding to a state in which the nouns (especially proper nouns) included in the utterance content have changed.
[0108] The feature change flag may be added to the feature database 122, for example, and as a result of the processing in step S6, combined data (plural sets) of the still image data ID, the determination feature, and the feature change flag may be accumulated and recorded in the feature database 122. Alternatively, combined data (plural sets) of the still image data ID and the feature change flag may be saved in a separate database or data file.
[0109] The element determination unit 117 reads out the combination data (note that there are multiple sets) of still image data IDs and image elements recorded in the element database 123 for the video data to be processed, and compares each image element contained in each still image of the still image data that are sequential in time series (note that the order in time series is specified, for example, by the still image data ID) with each other (step S7).
[0110] The element determination unit 117 compares each image element contained in each still image of the still image data that precedes and follows the still image data in the time series, for which the feature determination unit 116 has compared the determination features.
[0111] Specifically, the element determination unit 117 compares at least one of the following image elements contained in each still image data item that is one after the other in the time series: person identifiers, subtitles (text), subtitle genres, music titles, audio genres, and location identifiers (specifically, image elements that have been detected and extracted in the processing of step S5 and included in the element database 123).
[0112] When the person identifiers of successive still image data in the time series are different (for example, at least some of them are different or do not match at all), the element determination unit 117 determines that the space between these successive still image data is a candidate for a scene change point, and assigns a person change flag to the later still image data in the time series (although it may also be the earlier still image data in the time series) of these successive still image data (specifically, for example, the value of the person change flag, which has an initial value of 0, is changed to 1).
[0113] When the subtitles (text) of successive still image data in the time series are different (for example, at least a portion is different or they do not match at all), the element determination unit 117 determines that the space between these successive still image data is a candidate for a scene change point, and assigns a subtitle change flag to the later still image data in the time series (although it may also be the earlier still image data in the time series) of these successive still image data (specifically, for example, it changes the value of the subtitle change flag, which has an initial value of 0, to 1).
[0114] When the genres of the subtitles (text) of successive still image data in the time series are different, the element determination unit 117 determines that the space between these successive still image data is a candidate for a scene change point, and assigns a subtitle genre change flag to the later still image data in the time series (although it may also be the earlier still image data in the time series) of these successive still image data (specifically, for example, it changes the value of the subtitle genre change flag, which has an initial value of 0, to 1).
[0115] When the music titles of the successive still image data in the time series are different, the element determination unit 117 determines that the space between these successive still image data is a candidate for a scene change point, and assigns a music change flag to the later still image data in the time series (or the earlier still image data in the time series) of these successive still image data (specifically, for example, the value of the music change flag, which has an initial value of 0, is changed to 1).
[0116] When the audio genres of the successive still image data in the time series are different, the element determination unit 117 determines the space between these successive still image data as a candidate scene change point, and assigns an audio genre change flag to the later still image data in the time series (although it may also be the earlier still image data in the time series) of these successive still image data (specifically, for example, it changes the value of the audio genre change flag, which has an initial value of 0, to 1).
[0117] When the location identifiers of successive still image data in the time series are different, the element determination unit 117 determines that the space between these successive still image data is a candidate for a scene change point, and assigns a location change flag to the later still image data in the time series (although it may also be the earlier still image data in the time series) of these successive still image data (specifically, for example, the value of the location change flag, which has an initial value of 0, is changed to 1).
[0118] Then, if at least one of the person change flag, caption change flag, caption genre change flag, music change flag, audio genre change flag, and location change flag is assigned to each still image data (specifically, for example, the flag value is 1), the element determination unit 117 determines that the still image data is a candidate for a scene change point and assigns an element change flag to the still image data (specifically, for example, the value of the element change flag, which has an initial value of 0, is changed to 1).
[0119] Among the person change flag, caption change flag, caption genre change flag, music change flag, audio genre change flag, and location change flag, the flags that are required for the element determination unit 117 to determine that a scene is a candidate for a scene change point may be set.
[0120] For example, if the caption genre change flag is set to a required flag, and each piece of still image data is assigned a caption genre change flag (specifically, for example, the value of the caption genre change flag is 1), and at least one of the person change flag, caption change flag, music change flag, audio genre change flag, and location change flag is assigned (specifically, for example, the flag value is 1), the still image data may be determined to be a candidate for a scene change point, and an element change flag may be assigned to the still image data (specifically, for example, the value of the element change flag, which has an initial value of 0, may be changed to 1).
[0121] In the above case, the image element includes the genre of the caption as a required image element.
[0122] In addition, if the audio genre change flag is set to a required flag, and the audio genre change flag is assigned to each still image data (specifically, for example, the value of the audio genre change flag is 1), and at least one of the person change flag, subtitle change flag, subtitle genre change flag, music change flag, and location change flag is assigned (specifically, for example, the flag value is 1), the still image data may be determined to be a candidate for a scene change point, and an element change flag may be assigned to the still image data (specifically, for example, the value of the element change flag, which has an initial value of 0, may be changed to 1).
[0123] In the above case, the image element includes the audio genre as a required image element.
[0124] That is, in order to achieve the above, at least one of the genre of the telop and the genre of the audio may be made an essential image element.
[0125] Furthermore, if at least one of a caption genre change flag and an audio genre change flag is assigned (specifically, for example, the flag value is 1), the still image data may be judged as a candidate for a scene change point, and an element change flag may be assigned to the still image data (specifically, for example, the value of the element change flag, which has an initial value of 0, may be changed to 1).
[0126] The person change flag, telop change flag, telop genre change flag, music change flag, audio genre change flag, location change flag, and element change flag may be added to, for example, element database 123. As a result of the processing of step S7, combined data (plural sets) of still image data ID, person identifier, telop (text), telop genre, music title, audio genre, location identifier, person change flag, telop change flag, telop genre change flag, music change flag, audio genre change flag, location change flag, and element change flag may be accumulated and recorded in element database 123. Alternatively, combined data (plural sets) of still image data ID, person change flag, telop change flag, telop genre change flag, music change flag, audio genre change flag, location change flag, and element change flag may be saved in a separate database or data file.
[0127] The switching point identification unit 118 determines that each still image data is a scene switching point when at least one of a feature change flag and an element change flag is assigned to the still image data (specifically, for example, at least one of the feature change flag value and the element change flag value is 1) (step S8).
[0128] When a caption genre change flag is set as a required flag for determining whether an image element is a candidate for a scene change point (prerequisite: the image element contains the caption genre) or an audio genre change flag is set (prerequisite: the image element contains the audio genre), the condition for the change point identification unit 118 to determine whether an image element is a scene change point is that at least the element change flag is assigned.
[0129] The change point identification unit 118 may change the way of determining whether or not a scene change point is present depending on the genre of the program. For example, if the genre of the program is drama, the still image data may be determined to be a scene change point if both a feature amount change flag and an element change flag are attached, or if the genre of the program is news, the still image data may be determined to be a scene change point if an element change flag is attached (i.e., the determination is based only on the element change flag).
[0130] By the process of step S8, all scene change points (specifically, still image data IDs corresponding to (immediately after) scene change points) are detected for the video data to be processed. As a result, the video data to be processed is divided into scenes.
[0131] If each still image data ID is associated with the elapsed time from the start of the original video data to be processed, the elapsed time from the start of each scene change point in the video data to be processed is grasped and acquired. Therefore, the elapsed time from the start and end of each scene in the video data to be processed is grasped and acquired.
[0132] The summary generation unit 119 generates a summary for each scene divided in the process of step S8 (step S9). Note that the process of step S9 is not essential to the present invention, and therefore the summary generation unit 119 is not an essential component of the present invention.
[0133] The scene in the process of step S9 is, in other words, a section from a change point of one scene to a change point of the next scene in time series.
[0134] Specifically, the scene in the processing of step S9 is, for example, an image corresponding to the section from the still image data ID corresponding to immediately after a certain scene change point identified in the processing of step S8 to the still image data ID immediately before the still image data ID corresponding to immediately after the next scene change point in the time series.
[0135] The scene in the processing of step S9 is an image corresponding to the section from the elapsed time at the start of a scene to the elapsed time at the end (or the still image data ID immediately before the still image data ID corresponding to the elapsed time at the end) based on the elapsed time from the start of the original video data to be processed at each scene change point, which is grasped and obtained when each still image data ID is associated with the elapsed time from the start of the original video data to be processed.
[0136] For example, the summary generation unit 119 may use the results of all or at least one of the following processes i to iv for each scene as a summary of that scene. Alternatively, a scene summary may be generated from the results of multiple processes i to iv using a generative AI such as Google Vartex AI or ChatGPT, an LLM (large-scale language model), or the like. Alternatively, the video itself may be summarized by an AI using a generative AI such as Google Vartex AI or ChatGPT, an LLM (large-scale language model), or the like.
[0137] i) Speech to text Using speech recognition technology (Speech to Text), sounds emitted within the scene are detected and converted into text.
[0138] ii) Converting subtitles into text Using OCR technology, subtitles displayed within a scene are detected and converted into text.
[0139] iii) Textualization of things Using image recognition technology, various things such as people and objects displayed in the scene are detected and listed (in other words, converted into text).
[0140] iv) Attribute assignment Using natural language processing, attributes are assigned to the scene in question based on the text data generated by the processes i to iii above.
[0141] To assign attributes, for example, natural language processing is first used to break down the text data into parts of speech (specifically, nouns, verbs, adjectives, adverbs, etc.), and then named entity recognition (NER) is used to assign attributes such as "person's name," "place name," and "organization name."
[0142] In addition, proper nouns such as people's names and organization names are linked to separately prepared master data that organizes various attributes related to people's names and organization names, and are linked to attribute information such as reading, affiliated agency, official website URL, stock code, and corporate number.
[0143] By having the master data associate the "regular name" with the "variant spelling name," spelling variations can be absorbed and the names can be consolidated.
[0144] For proper nouns such as people's names and organization names, relevant attributes are assigned using generative AI such as ChatGPT or LLM (large-scale language model). For example, for people's names, reading, date of birth, affiliation, official website, and social media URLs may be assigned, and for organization names, information such as reading, establishment date, parent company, subsidiary, and other affiliated companies, official website, and social media URLs may be automatically collected and assigned.
[0145] (Action and effect) According to the video processing server 1 and video processing program of the embodiment, when at least one of a feature change flag and an element change flag is assigned, it is determined to be a scene change point, making it possible to accurately detect scene changes.
[0146] Furthermore, by performing the scene division process (step S8) and summary generation process (step S9), metadata for videos distributed in the world is organized, and by using the scene-by-scene video metadata, more detailed effectiveness measurement, trend analysis, and other processes can be performed. By performing the scene division process (step S8) and summary generation process (step S9), and by using the text obtained as a result of these processes to self-learn the annotated data of the AI, a cycle of automatic accuracy improvement can be created. Furthermore, this not only promises accuracy improvement through the automatic self-learning of the AI, but also makes it possible to utilize the data as annotation data for other AI solutions.
[0147] The above describes the embodiments of the present invention, but the specific configuration of the present invention is not limited to the above embodiments, and the present invention also includes forms in which modifications and changes are made to the above embodiments within the scope of the gist of the present invention.
[0148] For example, in the above embodiment, the video database 21 is stored in the cloud server 2, but the video database 21 may be stored in the storage unit 12 of the video processing server 1. Also, in the above embodiment, the still image database 121, the feature database 122, and the element database 123 are stored in the storage unit 12 of the video processing server 1, but at least one of the still image database, the feature database 122, and the element database 123 may be stored in the cloud server 2. [Explanation of symbols]
[0149] 1. Video processing server (video processing device, information processing device) 11 Control section 111 Instruction Reception Department 112 Data Acquisition Unit 113 Still image generation section 114 Feature Calculation Unit 115 Element Extraction Unit 116 Feature determination unit 117 Element judgment part 118 Switching point identification unit 119 Summary generator 12 Storage section 121 Still Image Database 122 Feature Database 123 Element Database 13 Input section 14 Communications Department 15 Drive device 19 Bus 2. Cloud Server 21 Video Database 9. Communication Networks 100 Video Processing System
Claims
1. means for acquiring a plurality of instantaneous image data at predetermined time intervals from the image data to be processed; means for calculating, for each of the plurality of pieces of video data at each moment, a feature amount of an image generated by the video data at that moment; means for detecting and extracting elements included in an image generated by the video data at each of the plurality of instants of time; means for determining, when the image features of the video data of the preceding and succeeding moments in time series are different, the interval between the video data of the preceding and succeeding moments as a candidate for a scene change point, and assigning a feature change flag to one of the video data of the preceding and succeeding moments; means for determining, when elements included in the images of the video data of successive moments in time series are different, the interval between the video data of successive moments as a candidate for a scene change point, and assigning an element change flag to one of the video data of successive moments; means for determining that the video data at the moment to which at least one of the feature amount change flag and the element change flag is assigned is a scene change point, The element included in the image is at least one of a person, a caption, a caption genre, music, an audio genre, and a location.
1. A video processing device comprising:
2. The feature of the image is any one of a color histogram, a shape feature, an edge feature, a HoG feature, and a CNN feature.
2. The video processing device according to claim 1,
3. The elements included in the image include at least one of the genre of the telop and the genre of the audio as essential elements.
2. The video processing device according to claim 1,
4. The video processing device further includes means for generating a summary for each section from a certain scene change point to a next scene change point in a time series among the scene change points determined to be the scene change points, the summary is a result of at least one of processes including converting speech into text, converting subtitles into text, converting things into text, and assigning attributes based on text data generated by converting speech into text, converting subtitles into text, and converting things into text; 2. The video processing device according to claim 1,
5. A process of acquiring video data of a plurality of moments at predetermined time intervals for the video data to be processed; A process of calculating, for each of the plurality of pieces of video data at each moment, a feature amount of an image generated by the video data at that moment; a process of detecting and extracting elements included in an image generated by the video data of each of the plurality of instants of time; a process of determining a point between the video data of the preceding and succeeding moments as a candidate for a scene change point when the image features of the respective video data of the preceding and succeeding moments in time series are different, and assigning a feature change flag to one of the video data of the preceding and succeeding moments; a process of determining a scene change point between the video data of the preceding and succeeding moments in time when the elements included in the images of the respective video data of the preceding and succeeding moments are different, and assigning an element change flag to one of the video data of the preceding and succeeding moments; determining that the video data at the moment to which at least one of the feature amount change flag and the element change flag is assigned is a scene change point; The element included in the image is at least one of a person, a caption, a caption genre, music, an audio genre, and a location. A video processing program characterized by:
6. The feature of the image is any one of a color histogram, a shape feature, an edge feature, a HoG feature, and a CNN feature.
6. The video processing program according to claim 5.
7. The elements included in the image include at least one of the genre of the telop and the genre of the audio as essential elements.
6. The video processing program according to claim 5.
8. further causing the information processing device to execute a process of generating a summary for each section from a certain scene change point to a next scene change point in a time series among the scene change points determined to be the scene change points; the summary is a result of at least one of processes including converting speech into text, converting subtitles into text, converting things into text, and assigning attributes based on text data generated by converting speech into text, converting subtitles into text, and converting things into text; 6. The video processing program according to claim 5.
Citation Information
Patent Citations
Method and apparatus for navigating content within a personal video recorder
JP2007502080A