Video processing device, video processing system, video processing method, and program

The video processing device classifies frames and generates text to convey the original content, addressing the loss of information in existing conversion technologies, thereby reducing storage and transmission needs while preserving content integrity.

JP2026055028APending Publication Date: 2026-03-30NEC CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-17
Publication Date
2026-03-30

AI Technical Summary

Technical Problem

Existing video conversion technologies, such as those described in Patent Document 1, may result in the loss of original video content, leading to users being unable to correctly grasp the intended content due to the replacement of representative frames with similar images, thus failing to convey the original video content effectively.

Method used

A video processing device and system that classifies image frames into distribution and non-distribution frames, generating text to describe the content between non-distribution frames, ensuring the original content is preserved and conveyed through distribution frames.

Benefits of technology

The solution maintains the integrity of the video content by reducing storage and transmission capacity without losing essential information, allowing users to accurately understand the original content through text descriptions of omitted frames.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026055028000001_ABST
    Figure 2026055028000001_ABST
Patent Text Reader

Abstract

To provide a video processing device, a video processing system, a video processing method, and a program for processing video in a way that does not impair the content of the video. [Solution] The video processing device according to this disclosure includes a classification means that classifies each of a plurality of image frames constituting a video into a distribution frame to be distributed and a non-distribution frame other than the distribution frame, and a generation means that generates text that describes the video content between two distribution frames based on the non-distribution frame contained between the two distribution frames.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a video processing apparatus, a video processing system, a video processing method, and a program.

Background Art

[0002] In recent years, various technologies for distributing video data have been proposed. For example, Patent Document 1 discloses a video conversion technique for converting the content of video data into document data, and a document conversion technique for re-converting the document data into the original video data. The video conversion technique according to Patent Document 1 generates a representative frame for a certain video segment, discovers an image similar to the representative frame from a database, and uses a reference value such as the URL (Uniform Resource Locator) of the image as part of the document data to convert the video data into document data.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, the document data obtained by the video conversion technique according to Patent Document 1 may not be able to sufficiently convey the original video content to the user. Specifically, since the video conversion is performed by replacing the representative frame of the video segment with a similar image in the image database, there is a possibility that a video essentially different from the original video content may be restored by the document conversion technique according to Patent Document 1. In this case, since the restored video data has lost some or all of the original video content, the user who views the video data may not be able to correctly grasp the original video content.

[0005] This disclosure is made to solve these problems and aims to provide a video processing device, a video processing system, a video processing method, and a program for processing video without damaging the video content. [Means for solving the problem]

[0006] The video processing device according to this disclosure includes a classification means for classifying each of a plurality of image frames constituting a video into a distribution frame to be distributed and a non-distribution frame other than the distribution frame, and a generation means for generating text that describes the video content between two distribution frames based on the non-distribution frame included between the two distribution frames.

[0007] The video processing system according to this disclosure comprises a video transmission device and a video playback device, wherein the video transmission device includes classification means for classifying each of a plurality of image frames constituting a video into distribution frames to be distributed and non-distribution frames other than the distribution frames, generation means for generating text describing the video content between two distribution frames based on the non-distribution frames included between the two distribution frames, and transmission means for transmitting the distribution frames and the text, and the video playback device includes playback means for playing back the distribution frames using the text received from the video transmission device.

[0008] The video processing method relating to this disclosure involves a computer classifying each of a plurality of image frames constituting a video into a distribution frame to be distributed and a non-distribution frame other than the distribution frame, and generating text that describes the video content between the two distribution frames based on the non-distribution frame contained between the two distribution frames.

[0009] The program relating to this disclosure causes a computer to perform the following processes: classify each of the multiple image frames constituting a video into a distribution frame to be distributed and a non-distribution frame other than the distribution frame; and generate text that describes the video content between the two distribution frames based on the non-distribution frame contained between the two distribution frames. [Effects of the Invention]

[0010] This disclosure provides a video processing device, a video processing system, a video processing method, and a program for processing video in a manner that does not impair the video content. [Brief explanation of the drawing]

[0011] [Figure 1] Figure 1 is a block diagram showing the configuration of the video processing device 1 according to this disclosure. [Figure 2] Figure 2 is a flowchart showing an example of the processing operation of the video processing device 1 according to this disclosure. [Figure 3] Figure 3 shows an example of video processing by the video processing device 1 according to this disclosure. [Figure 4] Figure 4 shows an example of video processing by the video processing device 1 according to this disclosure. [Figure 5] Figure 5 is a block diagram showing the configuration of the video processing device 1 according to this disclosure. [Figure 6] Figure 6 is a flowchart showing an example of the processing operation of the video processing device 1 according to this disclosure. [Figure 7] Figure 7 is a block diagram showing the configuration of the video processing system 30 related to this disclosure. [Figure 8] Figure 8 is a flowchart showing an example of the processing operation of the video processing system 30 according to this disclosure. [Figure 9] Figure 9 shows an example of the hardware configuration of the video processing device 40 according to this disclosure. [Modes for carrying out the invention]

[0012] First, I will explain in detail the problems that the technology disclosed herein aims to solve. In recent years, various video processing technologies have been used not only for video conversion technology related to Patent Document 1, but also for saving recorded videos. A typical video processing technology is the technology that encodes the video to be transmitted using the MPEG (Moving Picture Experts Group) standard, etc. In MPEG encoding, the number of image frames encoded as still images is suppressed by encoding the difference information from the previous frame and motion vectors. In addition, the MPEG standard also utilizes a technology that encodes the difference information between adjacent pixels within the same frame.

[0013] Furthermore, video processing methods using variable frame rate (VFR) and variable bitrate (VBR) technologies are also employed. VFR and VBR are technologies that dynamically adjust the frame rate and bitrate according to the video content. When there are sections in a video with little variation, VFR or VBR can be used to reduce the frame rate and bitrate in those sections. On the other hand, for sections in a video with a lot of variation, the video can be recorded and saved at a high frame rate and bitrate.

[0014] Furthermore, encoding techniques based on Region of Interest (ROI) are also used. ROI refers to the region of a video that is of high interest to the user. ROI encoding distinguishes objects in a video into ROI and non-ROI objects, assigning more encoding to ROI objects and less to non-ROI objects, thereby compressing the video according to the importance of each object. ROI encoding allows for video storage while maintaining high image quality of the ROI.

[0015] According to the above video processing technology, it is possible to process a video while maintaining a high image quality of the target video. On the other hand, in recent years, the need to reduce the video capacity has been increasing even more. When storing a video shot at a high resolution in a storage device such as a hard disk drive (HDD) or a solid state drive (SSD), such video data compresses the free capacity of these storage devices. In addition, when transmitting and receiving such video data via an Internet line, a large transmission capacity is required.

[0016] In the video processing technology as described above, it is not possible to significantly reduce the storage capacity of the video. In contrast, the video conversion technology according to Patent Document 1 can reduce the storage capacity by converting video data into document data.

[0017] However, in the document data obtained by the video conversion technology according to Patent Document 1, there is a possibility that the original video content cannot be sufficiently conveyed to the user. Specifically, since the video conversion is performed by replacing the representative frame of the video section with a similar image in the image database, there is a possibility that a video that is essentially different from the original video content may be restored by the document conversion technology according to Patent Document 1. In this case, since the restored video data has lost some or all of the original video content, the user who views the video data cannot correctly grasp the original video content.

[0018] Each embodiment according to the present disclosure described below contributes to the solution of the above problems.

[0019] (Embodiment 1) Embodiment 1 of the present disclosure will be described below with reference to the drawings. Figure 1 is a block diagram showing the configuration of the video processing device 1 according to the present disclosure. The video processing device 1 is, for example, a server, cloud, personal computer (PC), smartphone, tablet terminal, or the like. The video processing device 1 also includes a classification unit 11 and a generation unit 12. The classification unit 11 and the generation unit 12 may be used as means for classifying information or data and means for generating information or data, respectively.

[0020] The video processed by the video processing device 1 includes video data consisting of multiple image frames. The video to be processed may be recorded at a fixed frame rate or at a variable frame rate. Furthermore, the video may or may not include audio data. In addition, the video may be unencoded or encoded. When processing compressed video, the video processing device 1 may process the video by decoding it. Hereafter, the video to be processed by the video processing device 1 will be referred to as the target video, and the image frames related to the target video will be referred to as the target image frames.

[0021] The classification unit 11 classifies each of the multiple target image frames that make up the target video into distribution frames and non-distribution frames. Distribution frames are the image frames among the target image frames that are to be saved, and non-distribution frames are the image frames among the target image frames that are not to be saved. In other words, the classification unit 11 classifies the multiple target image frames into image frames to be saved and image frames that are not to be saved. Depending on the use of the distribution frames, distribution frames can also be described as image frames that are to be distributed, and non-distribution frames can also be described as image frames that are not to be distributed. Alternatively, distribution frames can be described as image frames that are to be output, and non-distribution frames can also be described as image frames that are not to be output. Several examples of the classification method used by the classification unit 11 are shown below.

[0022] The classification unit 11 may classify the target image frames into distribution frames and non-distribution frames at predetermined intervals. That is, the classification unit 11 may classify the target image frames arranged in playback order into distribution frames and non-distribution frames at equal intervals. Here, the playback order is determined, for example, based on the time code assigned to the target image frames. For example, the classification unit 11 may divide the target image frames arranged in playback order into groups of three frames, designating the first frame as a distribution frame and the following two frames as non-distribution frames. Thereafter, when a frame is said to follow another frame, it means that it follows in the playback order.

[0023] In the classification method described above, if the frame rate of the target video is 30fps (frames per second), the classification unit 11 selects the first image frame as a distribution frame and the following two frames as non-distribution frames for any given 1-second frame of the target image. Then, the classification unit 11 selects the next frame as a distribution frame. In this way, the classification unit 11 classifies 10 frames as distribution frames and 20 frames as non-distribution frames at equal intervals. In other words, if the playback time of the distribution frames relative to the reference clock does not change before and after classification, the frame rate of the distribution frames after classification will be 10fps. In this embodiment, it is assumed that the playback time of the distribution frames relative to the reference clock does not change before and after classification.

[0024] Furthermore, the classification unit 11 may classify the distributed frames into distributed frames and non-distributed frames so that the frame rate of the distributed frames is a predetermined value. If the target video is recorded at a fixed frame rate and the set predetermined value is less than that frame rate, the classification unit 11 may classify the target video so that its frame rate is a predetermined value.

[0025] For example, if the frame rate of the target video is 30fps and a predetermined value is set to 10fps, the classification unit 11 will classify the distributed frames so that their frame rate is 10fps. When classifying the distributed frames so that they are at equal intervals, the classification unit 11 classifies the two frames immediately following a given distributed frame as non-distributed frames. The classification unit 11 does not have to classify the distributed frames at equal intervals. Also, if the frame rate of the target video is lower than the set predetermined value, the classification unit 11 may classify all of the target image frames as distributed frames.

[0026] In the above classification method, if the target video is recorded at a variable frame rate, the classification unit 11 may determine the distribution frames in accordance with the set frame rate. In other words, the classification unit 11 may classify the distribution frames into distribution frames and non-distribution frames such that the frame rate of the distribution frames is less than or equal to a predetermined value. That is, if the frame rate in a certain section is greater than a set predetermined value, the classification unit 11 may determine the distribution frames such that the frame rate in that section becomes the predetermined value. On the other hand, if the frame rate in another section is less than a set predetermined value, the classification unit 11 may treat all of the target image frames in that section as distribution frames.

[0027] For example, consider a case where a target video recorded with a variable frame rate has a frame rate of 30fps in one section and a frame rate of 10fps in another section, and the predetermined value is set to 15fps. In this case, the classification unit 11 determines the distribution frames for the 30fps section so that the frame rate of the distribution frames is 15fps. That is, the classification unit 11 classifies 15 frames as distribution frames and 15 frames as non-distribution frames in that section. Also, for the 10fps section, since it is below the predetermined value of 15fps, the classification unit 11 classifies all target image frames in that section as distribution frames.

[0028] Furthermore, the classification unit 11 may determine the distribution frames such that the frame rate of the distribution frames is less than the frame rate of the target video by a predetermined amount. In other words, the classification unit 11 may determine the distribution frames by performing a relative classification with respect to the frame rate of the target video. For example, if the target video is recorded at a fixed frame rate and that frame rate is 30fps, the classification unit 11 may determine the distribution frames such that the frame rate of the distribution frames is uniformly 5fps less than that of the target video. That is, in this example, the classification unit 11 may determine the distribution frames such that the frame rate of the distribution frames is 25fps.

[0029] In the classification method described above, even if the target video is recorded at a variable frame rate, the classification unit 11 may determine the distribution frames so that the frame rate is uniformly reduced by a predetermined value. In this case, the classification unit 11 may set a lower limit on the frame rate of the distribution frames. For example, consider a case where the classification unit 11 determines the distribution frames so that the frame rate is uniformly reduced by 5 fps, and the lower limit on the frame rate is 5 fps. If the frame rate of a certain section is 30 fps, the classification unit 11 determines the distribution frames so that the frame rate of the distribution frames related to that section is 25 fps. On the other hand, if the frame rate of another section is 7 fps, the classification unit 11 determines the distribution frames so that the frame rate of the distribution frames related to that section is the lower limit of 5 fps.

[0030] Furthermore, the classification unit 11 may classify target image frames into distribution frames and non-distribution frames by comparing the target image frames and calculating the difference between frames and the motion vector relative to the reference frame. In this case, the classification unit 11 may calculate the amount of change between frames from the difference between frames and the motion vector, and if the amount of change is greater than a predetermined value, the frame may be designated as a distribution frame. In this case, the target image frames corresponding to areas in the target video where the change in the object is large are designated as distribution frames. Alternatively, the classification unit 11 may designate a frame as a distribution frame if the amount of change between frames is smaller than a predetermined value. In this case, the target image frames corresponding to areas in the target video where the change in the object is small are designated as distribution frames.

[0031] In the classification method described above, the classification unit 11 may compare a reference frame with a subsequent frame adjacent to the reference frame when classifying. Here, frames being adjacent means that their playback order is adjacent. For example, if the amount of change between the reference frame and the subsequent frame adjacent to the reference frame is greater than a predetermined value, the classification unit 11 may either classify both frames as distribution frames or both frames as non-distribution frames. Alternatively, the classification unit 11 may classify the reference frame as a distribution frame and the subsequent frame as a non-distribution frame.

[0032] The classification unit 11 may also compare a reference frame with a subsequent frame that is a predetermined number of frames away from the reference frame. If the amount of change between the reference frame and the subsequent frame is greater than a predetermined value, the classification unit 11 may designate both frames and all frames in between as distribution frames, or it may designate both frames and all frames in between as non-distribution frames. Alternatively, the classification unit 11 may designate both frames as distribution frames and the frames in between as non-distribution frames.

[0033] In this case, the classification unit 11 may then use the said subsequent frame as the next reference frame, and use a subsequent frame that is a predetermined number of frames away from the said next reference frame as the next subsequent frame, calculate the amount of change in the same manner as above, and perform the classification. In this way, the classification unit 11 may classify all target image frames into distributed frames and non-distributed frames.

[0034] Furthermore, the classification unit 11 may classify the target image frames into distribution frames and non-distribution frames using a trained model. That is, the classification unit 11 may classify the target image frames by using a trained model that accepts the target image frames as input data and outputs the classification results of distribution frames and non-distribution frames as output data. Here, the trained model may be one that has been machine-trained using pairs of target image frames and classification results as training data.

[0035] Specifically, the trained model may use image frames from videos similar to the target video, along with classification results obtained by classifying them according to certain rules, as training data. In this case, the trained model is generated by supervised learning, which uses a known machine learning algorithm such as a neural network to learn the relationship between the image frames of the similar videos and the classification results.

[0036] Next, the generation unit 12 will be described. The generation unit 12 generates text that describes the video content between two distribution frames, based on the non-distribution frames included between the two distribution frames. The generation unit 12 may generate text that describes the video content taking into account both the two distribution frames and the non-distribution frames included between them, or it may generate text that describes the video content taking into account only the non-distribution frames.

[0037] The two distribution frames may or may not be adjacent. That is, the generation unit 12 may generate text describing the video content between two adjacent distribution frames based on the non-distribution frames included between those two adjacent distribution frames. Alternatively, the generation unit 12 may generate text describing the video content between two non-adjacent distribution frames based on the non-distribution frames included between those two non-adjacent distribution frames and other distribution frames.

[0038] The video content that the generation unit 12 uses to generate text is, for example, a change that occurs in an object within the video between two streaming frames. An object is any object present in the target video. An object includes not only tangible things such as humans, animals, and plants, but also intangible things such as light. In other words, the video content could be, for example, an animal appearing from outside the frame into the frame, or a flashing of light occurring between two streaming frames.

[0039] The video content may include changes due to camera work. For example, the video content may include panning (horizontal movement), tilting (vertical movement), zooming, and rotation. The video content may also include changes resulting from changes in camera settings. For example, the video content may include changes resulting from adjusting focus, exposure, or white balance. Furthermore, the video content may include changes resulting from video editing between two streaming frames. For example, the video content may include changes resulting from scene cuts.

[0040] The text generated by the generation unit 12 is information for outputting the video content between two distribution frames to the outside of the video processing device 1. For example, the text may be used to convey the video content to the user. Here, the text may be a single sentence or composed of multiple sentences. Furthermore, if the text uses Japanese, it may consist of a subject and a predicate. In addition, the text may consist only of words.

[0041] The text generated by the generation unit 12 may be used to transmit video content to another device outside the video processing device 1. In this case, the text does not need to be decipherable by a user, as long as it can transmit the video content to the other device. That is, the text may be a string of characters written in a predetermined programming language.

[0042] Here, a device located outside the video processing device 1 is, for example, a subtitle information generation device. The subtitle information generation device generates subtitle information for displaying the video content together with the subsequent distribution frame among two distribution frames. In other words, the text generated by the generation unit 12 may be text for generating subtitle information to be displayed together with the said subsequent distribution frame.

[0043] Furthermore, the device located outside the video processing device 1 may be an audio information generation device. The audio information generation device generates audio information for outputting the video content as audio, for example, along with the subsequent distribution frame of two distribution frames. If the target video has audio, the audio information generation device is a device that generates secondary audio related to the video content.

[0044] The generation unit 12 may generate text that explains the video content between two streamed frames using a trained model. That is, the generation unit 12 may generate text by using a trained model that accepts two streamed frames and a non-streamed frame as input data and outputs text as output data. Here, the trained model may be one that has been machine-trained using pairs of two streamed frames, a non-streamed frame between them, and text as training data. Specifically, the generation unit 12 may generate text using a VLM (Vision-Language Model).

[0045] The following are several examples of conditions under which the generation unit 12 generates text. First, the generation unit 12 may always generate text if there is a non-delivery frame between two delivery frames. In other words, if there is no non-delivery frame between two delivery frames, the generation unit 12 does not need to generate text. Also, the generation unit 12 may generate text if the number of frames between two delivery frames is greater than or equal to a predetermined value.

[0046] Furthermore, if the target video is recorded at a fixed frame rate, the generation unit 12 may generate text if the frame rate between the two target distribution frames is less than or equal to a predetermined value relative to the frame rate of the target image frame. In other words, the generation unit 12 may generate text if the difference between the frame rate of the target image frame and the frame rate between the two target distribution frames is greater than or equal to a predetermined value.

[0047] If the target video is recorded at a fixed frame rate, and as a result of classification by the classification unit 11, the frame rate of the distribution frames changes relative to the frame rate of the target video, the generation unit 12 may generate text including the content of the non-distribution frames, i.e., missing frames.

[0048] If the two distribution frames targeted by the generation unit 12 are not adjacent, the frame rates between the two distribution frames may differ depending on the classification performed by the classification unit 11. In this case, the generation unit 12 may compare the frame rate of the target image frame with the average value of the frame rates between the two distribution frames, and generate text if the difference is greater than or equal to a predetermined value.

[0049] If the target video is recorded at a variable frame rate, and the distributed frames are also at a variable frame rate, the generation unit 12 may generate text only when the frame rate between two distributed frames is less than or equal to a predetermined value. In other words, the generation unit 12 does not need to generate text if the frame rate between two distributed frames is greater than the predetermined value. Here, the predetermined value is assumed to be a fixed value (absolute value). For example, consider the case where the predetermined value is 20fps. If the distributed frames are classified so that the frame rate is a maximum of 30fps and a minimum of 10fps, the generation unit 12 will not generate text for the interval where the frame rate between two distributed frames is 30fps. On the other hand, the generation unit 12 will generate text for the interval where the frame rate between two distributed frames is 10fps.

[0050] Even if the target video is recorded at a variable frame rate, if the classification by the classification unit 11 results in a change in the frame rate of the distributed frames in a certain section relative to the frame rate of the target video, the generation unit 12 may generate text including the content of the non-distributed frames, i.e., the missing frames.

[0051] In the above case, the generation unit 12 may set the number of non-delivered frames included between the two delivered frames as a predetermined value. That is, the generation unit 12 may generate text when the number of non-delivered frames included between the two delivered frames exceeds the predetermined value.

[0052] On the other hand, the generation unit 12 may set a predetermined value for the difference between the frame rate between two distribution frames and the frame rate of the corresponding target image frame. In other words, the predetermined value may be a relative value. That is, the generation unit 12 may generate text if the frame rate between two distribution frames is lower than or equal to the predetermined value compared to the frame rate of the corresponding target image frame. For example, consider the case where the predetermined value is 10fps. If the frame rate of the target image frame in a certain section is 30fps and the frame rate of the distribution frame in that section is 25fps, the difference is 5fps, which is lower than the predetermined value of 10fps. In this case, the generation unit 12 does not generate text for that section.

[0053] Next, an example of the processing flow by the video processing device 1 will be described. Figure 2 is a flowchart showing an example of the processing operation of the video processing device 1 according to this disclosure.

[0054] First, the classification unit 11 classifies each of the target image frames into a distribution frame and a non-distribution frame (S11). Here, the classification unit 11 may perform the classification process based on instructions from the user, or it may perform the classification process by receiving a request signal from an external device.

[0055] Next, the generation unit 12 generates text describing the video content between the two distribution frames based on the non-distribution frames included between the two distribution frames (S12). Here, the generation unit 12 may generate text after the classification unit 11 has performed classification processing on all of the target image frames. Alternatively, the generation unit 12 may generate text for the classified distribution frames even if the classification unit 11 has performed classification processing on only a portion of the target image frames.

[0056] An example of video processing by the video processing device 1 according to this disclosure will be described. Figure 3 is a diagram showing an example of video processing by the video processing device 1 according to this disclosure. In Figure 3, there are a total of 4 target image frames. The frames arranged in playback order are referred to as frame #1, frame #2, frame #3, and frame #4 in order. Here, frame #1 is the image frame at the reference clock time during video playback, and the playback time is set to t=0. The playback times of frame #2, frame #3, and frame #4, which follow frame #1, are set to t=1, t=2, and t=3, respectively. The arrows in the figure represent the playback order.

[0057] Figure 3 shows the movements of the cow and horse, which are objects in the video, from frame #1 to frame #4. In frame #1, the cow is standing on the left side of the image frame and is facing left. From frame #2 to frame #3, the horse walks in from outside the right side of the image frame and enters the frame. The horse is facing left of the image frame. In frame #4, the entire horse is within the image frame and stops to the right of the cow, that is, behind the cow.

[0058] The classification unit 11 classifies the target image frames into distribution frames and non-distribution frames using a predetermined method. As a result of the classification, the classification unit 11 classifies frames #1 and #4 as distribution frames and frames #2 and #3 as non-distribution frames. In this case, frames #1 and #4 become image frames to be saved, while frames #2 and #3 become image frames not to be saved.

[0059] The generation unit 12 generates text describing the video content between two adjacent distributed frames, frame #1 and frame #4, based on the non-distributed frames, frame #2 and frame #3, which are located between the two adjacent distributed frames, frame #1 and frame #4. Here, the generation unit 12 generates text as information to be used to convey the video content to the user. Furthermore, the generation unit 12 generates the text, "The horse appeared walking from behind the cow."

[0060] Figure 4 is a diagram illustrating an example of video processing by the video processing device 1 according to this disclosure. Specifically, Figure 4 is an example of a video saved when video processing by the video processing device 1 is applied to the video in Figure 3. Here, the text generated by the generation unit 12 is displayed as a subtitle in the subsequent distribution frame. As shown in Figure 4, the video processing device 1 saves only the distribution frames, frame #1 and frame #4. In addition, frame #4 displays the text generated by the generation unit 12, "The horse appeared walking from behind the cow."

[0061] Thus, the video processing device 1 according to this disclosure can process video without damaging the video content. In other words, since the classification unit 11 does not save non-distributed frames, it cannot grasp the video content between two distributed frames, and the user cannot correctly grasp the original video content. However, by having the generation unit 12 generate text that explains the video content, the video processing device 1 according to this disclosure can process video without damaging the original video content.

[0062] Furthermore, by using encoding methods such as MPEG and related technologies such as VFR, it is possible to process videos while maintaining high image quality. However, these technologies cannot significantly reduce the storage capacity of the videos. On the other hand, the video processing device 1 according to this disclosure can reduce the storage capacity of videos by thinning out non-distributed frames instead of saving them. In addition, it is possible to understand the content of the thinned-out image frames by transcribing the events of those image frames into text. As a result, the video processing device 1 according to this disclosure can process videos without damaging the video content while reducing the storage capacity of the videos.

[0063] (Embodiment 2) Next, a video processing device 1a according to this embodiment will be described. Figure 5 is a block diagram showing the configuration of the video processing device 1a according to this disclosure. The video processing device 1a comprises a classification unit 11, a generation unit 12, and an output unit 13. That is, the video processing device 1a according to this embodiment is the video processing device 1 according to Embodiment 1 with an output unit 13 added to its configuration. The output unit 13 may be used as a means for outputting information or data. The classification unit 11 and the generation unit 12 according to this embodiment are the same as those in Embodiment 1, so their description will be omitted.

[0064] The output unit 13 outputs the distribution frames classified by the classification unit 11 and the text generated by the generation unit 12. The output unit 13 may output only the distribution frames, only the generated text, or both. Furthermore, the output unit 13 may output the distribution frames and generated text simultaneously, consecutively, or at different timings. In addition, the output unit 13 may also output the audio information of the target video corresponding to the distribution frames. Hereafter, the content output by the output unit 13 will be referred to as the output content. The output unit 13 will output both the distribution frames and the generated text.

[0065] The output unit 13 may also be used as a transmission unit to transmit output content to another device. In this case, the output unit 13 may transmit output content by performing data communication with the other device via an internet connection or the like. The output unit 13 may also enable data communication via short-range wireless communication such as Wi-Fi® or Bluetooth®. Furthermore, the output unit 13 may perform data communication by connecting to the other device via a wired connection. The other device may be a terminal used by the user that is equipped with a means for playing back the output content. In this case, the other device may be, for example, a PC, smartphone, or tablet terminal. The other device may also be another server or the like for transmitting output content to the terminal used by the user.

[0066] Furthermore, the output unit 13 may be used as a playback means for playing back the output content in the video processing device 1a. In this case, the output unit 13 may have a display means such as a liquid crystal panel or organic electroluminescence. The output unit 13 may also have an audio output means such as a speaker or earphones, or it may have a connection terminal that can be connected to such audio output means.

[0067] The output unit 13 may convert the distribution frame and generated text into a predetermined playback format and output them, or it may output the distribution frame and generated text in their original format. In this case, when converting to the predetermined playback format, the output unit 13 may correct the playback time of the distribution frame with respect to the reference clock.

[0068] The output unit 13 may output the generated text as subtitle information. Here, subtitle information refers to subtitle information for including the generated text as subtitles in the image of the distribution frame. The subtitle information output by the output unit 13 may also be metadata for the image of the distribution frame. The metadata shall include the playback time of each text relative to the reference clock, along with the generated text. The subtitle information may be displayed as an overlay on the image of the corresponding distribution frame during playback, or it may be displayed outside the frame of the corresponding distribution frame. Alternatively, the subtitle information output by the output unit 13 may have the subtitles embedded in the image of the corresponding distribution frame. In other words, the output unit 13 may generate multiple image frames in which the distribution frame and the corresponding subtitle information are integrated, and output these.

[0069] The output unit 13 may output the generated text so that it is displayed as subtitle information along with the subsequent distribution frame (the subsequent playback frame) of the two corresponding distribution frames. Alternatively, the output unit 13 may output the text so that it is displayed along with the preceding distribution frame (the preceding playback frame).

[0070] Here, the output unit 13 may output content such that it displays subtitle information only for the duration of the subsequent or preceding playback frame. Alternatively, the output unit 13 may output content such that it displays the subtitle information continuously for a certain period of time or a certain number of frames from the subsequent or preceding playback frame. Specifically, the output unit 13 may determine the duration or number of frames to be continuously displayed from the subsequent or preceding playback frame according to the length (character length) of the subtitle to be displayed. For example, the output unit 13 may determine the display duration or number of frames in proportion to the length of the subtitle to be displayed.

[0071] The output unit 13 may output the generated text in a manner that allows it to be played back as audio. In other words, the output unit 13 may output the generated text as audio data. If the target video has audio, the output unit 13 will output the generated text in a manner that allows it to be played back as a secondary audio track.

[0072] The output unit 13 may output the output content such that the generated text is played along with the subsequent playback frame. Alternatively, the output unit 13 may output the output content such that the audio is played along with the preceding playback frame. Here, the output unit 13 may modify the playback time of the distribution frame so that playback of the distribution frame is stopped while the audio corresponding to the generated text is being played. That is, if the audio corresponding to the subsequent playback frame or preceding playback frame is played simultaneously with the display of the said playback frame, the output unit 13 may output the output content such that it continues to display the subsequent playback frame or preceding playback frame until the playback of the said audio is finished.

[0073] Next, an example of the processing flow of the video processing device 1a will be described. Figure 6 is a flowchart showing an example of the processing operation of the video processing device 1a according to this disclosure. Note that explanations that overlap with the processing operation of the video processing device 1 according to Embodiment 1 will be omitted as appropriate.

[0074] First, the classification unit 11 classifies each of the target image frames constituting the target video into a distribution frame and a non-distribution frame (S21). Next, the generation unit 12 generates a text describing the video content between the two distribution frames based on the non-distribution frame contained between the two distribution frames (S22). After that, the output unit 13 outputs the distribution frames and the text (S23). Here, the output unit 13 may output the output content only after the classification unit 11 has performed the classification process on all of the target image frames and the generation unit 12 has generated all the text. Alternatively, the output unit 13 may output the classified distribution frames even if the classification unit 11 has only performed the classification process on some of the target image frames. Furthermore, the output unit 13 may output the text even if the generation unit 12 has only generated text for some of the target distribution frames.

[0075] Thus, the video processing device 1a according to this embodiment can correctly transmit the original video content to the user by providing an output unit 13 to the video processing device 1 according to Embodiment 1. In other words, the video processing device 1a according to this embodiment can correctly transmit the original video content related to the video processed by the video processing device 1a to the user by outputting the text generated by the generation unit 12 so that it can be played back as subtitles, audio, etc.

[0076] Furthermore, when the output unit 13 transmits the output content to another device, the video processing device 1a according to this embodiment can reduce the transmission capacity of the video. Here, a certain degree of reduction in transmission capacity can be expected by using encoding according to the MPEG standard or related technologies such as VFR. However, even when these technologies are used, the transmission of image frames remains the same, so there is a limit to the effect of reducing transmission capacity. However, the video processing device 1a according to this embodiment can reduce the video capacity compared to the related technologies mentioned above because it decimates non-distribution frames by not including them in the transmission target. As a result, the video processing device 1a according to this embodiment can further reduce the transmission capacity compared to when only related technologies are used.

[0077] (Embodiment 3) Next, a video processing system 30 according to this embodiment will be described. Figure 7 is a block diagram showing the configuration of the video processing system 30 according to this disclosure. The video processing system 30 comprises a video transmission device 10 and a video playback device 20. The video transmission device 10 is a device that transmits the video processed by the video transmission device 10. Specifically, the video transmission device 10 is a device in which the output unit 13 of the video processing device 1a according to Embodiment 2 is replaced by a transmission unit 14. The transmission unit 14 may be used as a means for transmitting information or data. In other words, the video transmission device 10 uses the output unit 13 of the video processing device 1a according to Embodiment 2 as a transmission means. The video transmission device 10 is, for example, a server or a cloud. The configuration of the video transmission device 10 is the same as that of the video processing device 1a according to Embodiment 2, so a description of its configuration will be omitted.

[0078] The video playback device 20 is a device that receives output content transmitted from the video transmission device 10 and plays the video related to that output content. The video playback device 20 may or may not permanently store the output content received from the video transmission device 10. The video playback device 20 may be a terminal used by the user. For example, the video playback device 20 may be a PC, smartphone, tablet, etc.

[0079] The video transmission device 10 and the video playback device 20 are capable of communicating data with each other. Data communication between the video transmission device 10 and the video playback device 20 may be achieved via an internet connection, or through short-range wireless communication such as Wi-Fi or Bluetooth. Alternatively, data communication may be achieved by connecting the video transmission device 10 and the video playback device 20 via a wired connection.

[0080] The video playback device 20 includes a playback unit 21. The playback unit 21 can be described as a means for playing back the output content received from the video transmission device 10. Specifically, the playback unit 21 can be described as a means for playing back distribution frames using generated text received from the video transmission device 10. The playback unit 21 may also have a display means such as a liquid crystal panel or organic electroluminescence. Furthermore, the playback unit 21 may have an audio output means such as a speaker or earphones, or it may have a connection terminal that can be connected to such audio output means.

[0081] Next, an example of the processing flow of the video processing system 30 will be described. Figure 8 is a flowchart showing an example of the processing operation of the video processing system 30 according to this disclosure. Regarding the processing operation related to the video transmission device 10, explanations that overlap with those of the video processing device 1 according to Embodiment 1 and the video processing device 1a according to Embodiment 2 will be omitted as appropriate.

[0082] First, the classification unit 11 of the video transmission device 10 classifies each of the target image frames constituting the target video into a distribution frame and a non-distribution frame (S31). Next, the generation unit 12 of the video transmission device 10 generates text that explains the video content between the two distribution frames, based on the non-distribution frame contained between the two distribution frames (S32). Then, the transmission unit 14 of the video transmission device 10 transmits the distribution frames and text to the video playback device 20 (S33). After that, the playback unit 21 of the video playback device 20 receives the distribution frames and text from the video transmission device 10 and plays the video (S34). Here, the playback unit 21 may play the video after receiving all the distribution frames and text, or it may play the video even if it has only received some of the distribution frames and text.

[0083] Thus, the video processing system 30 according to this disclosure can accurately transmit the original video content of the video processed by the video transmission device 10 to the user while reducing transmission capacity. In other words, the video processing system 30 according to this disclosure can realize a series of processes from classifying image frames related to the target video, generating text that explains the video content, to playing these back.

[0084] (Example hardware configuration) Figure 9 shows an example of the hardware configuration of a video processing device 40 according to the present disclosure. In Figure 8, the video processing device 40 has a processor 41 and a memory 42. The processor 41 may be, for example, a microprocessor, an MPU (Micro Processing Unit), or a CPU (Central Processing Unit). The processor 41 may include multiple processors. The memory 42 is composed of a combination of volatile memory and non-volatile memory. The memory 42 may include storage located away from the processor 41. In this case, the processor 41 may access the memory 42 via an I / O (Input / Output) interface, which is not shown.

[0085] In the above example, the program can be stored and provided to the computer using various types of non-transitory computer-readable medium. Non-transitory computer-readable medium includes various types of tangible storage medium. Examples of non-transitory computer-readable medium include magnetic storage media (e.g., magneto-optical disks), CD-ROMs, CD-Rs, CD-R / Ws, and semiconductor memory (e.g., mask ROMs, PROMs (Programmable ROMs), EPROMs (Erasable PROMs), flash ROMs, RAMs). Alternatively, the program may be provided to the computer using various types of transient computer-readable medium. Examples of transient computer-readable medium include electrical signals, optical signals, and electromagnetic waves. Transitory computer-readable medium can supply the program to the computer via wired communication channels such as electric wires and optical fibers, or via wireless communication channels. Computers include various information processing devices such as PCs, servers, CPUs, MPUs, FPGAs (Field Programmable Gate Arrays), and ASICs (Application Specific Integrated Circuits).

[0086] Although the present disclosure has been described above with reference to embodiments, the present disclosure is not limited to the embodiments described above. Various modifications to the structure and details of the present disclosure can be made as can be understood by those skilled in the art within the scope of the present disclosure. Furthermore, each embodiment can be combined with other embodiments as appropriate.

[0087] Each drawing is merely illustrative to illustrate one or more embodiments. Each drawing may be associated with one or more other embodiments rather than with only one specific embodiment. As those skilled in the art will understand, various features or steps described with reference to any one drawing can be combined with features or steps shown in one or more other drawings, for example, to create embodiments not explicitly shown or described. Not all features or steps shown in any one drawing to illustrate an exemplary embodiment are necessarily required, and some features or steps may be omitted. The order of steps shown in any of the drawings may be changed as appropriate.

[0088] Some or all of the above embodiments may also be described as follows, but are not limited to the following: (Note 1) A classification means for classifying each of the multiple image frames that make up a video into a distribution frame to be distributed and a non-distribution frame other than the distribution frame, A generation means that generates text describing the video content between the two distribution frames based on the non-distribution frame contained between the two distribution frames, A video processing device equipped with the following features. (Note 2) The generation means generates text describing the video content between two adjacent distribution frames, based on the non-distribution frames that are included between two adjacent distribution frames in the playback order among the plurality of distribution frames. The video processing device described in Appendix 1. (Note 3) The generation means generates text as video content that describes the changes that occurred in the objects within the video between the two distribution frames. The video processing device described in Appendix 1 or 2. (Note 4) The generation means generates the text as information used to convey the video content to the user. A video processing device as described in any one of the items from Appendix 1 to Appendix 3. (Note 5) The generation means generates subtitle information as text to be displayed together with the subsequent playback frame, which is in a later playback order among the two distribution frames. A video processing device as described in any one of the items from Appendix 1 to Appendix 4. (Note 6) The classification means determines the distribution frame in accordance with the set frame rate, The generation means generates the text when the frame rate is less than or equal to a predetermined value. A video processing device as described in any one of the items from Appendix 1 to Appendix 5. (Note 7) The generation means generates the text when the frame rate of the two distribution frames in a specific section becomes lower than or equal to a predetermined value compared to the frame rate of the plurality of image frames in that section. A video processing device as described in any one of the items from Appendix 1 to Appendix 5. (Note 8) The classification means determines the distribution frame in accordance with the set frame rate, The generation means generates the text when the number of non-delivery frames included between the two determined delivery frames exceeds a predetermined number. A video processing device as described in any one of the items from Appendix 1 to Appendix 5. (Note 9) The video processing device further comprises output means for outputting the distribution frame and the text. A video processing device as described in any one of the items from Appendix 1 to Appendix 8. (Note 10) The output means outputs the subsequent playback frame and the text so that they are displayed as subtitle information together with the subsequent playback frame, which is played later in the playback order of the two distribution frames. The video processing device described in Appendix 9. (Note 11) The output means outputs the subsequent playback frame and the text so as to output the text as audio along with the display of the subsequent playback frame, which is in a later playback order among the two distribution frames. The video processing apparatus described in Appendix 9 or 10. (Note 12) The classification means accepts the image frame as input data, outputs the classification result of either the distributed frame or the non-distributed frame as output data, and performs the classification using a trained model that has been machine-trained using the pair of the image frame and the classification result as training data. A video processing device as described in any one of the items from Appendix 1 to Appendix 11. (Note 13) The generation means receives the two distribution frames and the non-distribution frame as input data, uses the text as output data, and generates the text using a trained model that has been trained on machine learning using the two distribution frames and the non-distribution frame contained between the two distribution frames and the text as training data. A video processing device as described in any one of the items from Appendix 1 to Appendix 12. (Note 14) It has a video transmission device and a video playback device, The aforementioned video transmission device is A classification means for classifying each of the multiple image frames that make up a video into a distribution frame to be distributed and a non-distribution frame other than the distribution frame, A generation means that generates text describing the video content between the two distribution frames based on the non-distribution frame contained between the two distribution frames, The system includes a transmission means for transmitting the distribution frame and the text, The aforementioned video playback device is The system includes a playback means for playing back the distribution frame using the text received from the video transmission device. Video processing system. (Note 15) Computers Each of the multiple image frames that make up the video is classified into a distribution frame to be distributed and a non-distribution frame other than the distribution frame. Based on the non-distributed frame contained between the two distribution frames, a text describing the video content between the two distribution frames is generated. Video processing methods. (Note 16) A process that classifies each of the multiple image frames that make up a video into a distribution frame to be distributed and a non-distribution frame other than the distribution frame, A process to generate text describing the video content between the two distribution frames, based on the non-distribution frame contained between the two distribution frames, A program that causes a computer to execute something.

[0089] Some or all of the elements (e.g., configuration and function) described in Appendices 2 to 13 that are dependent on Appendice 1 may also be dependent on Appendices 14, 15, and 16 in the same manner as those described in Appendices 2 to 13. Some or all of the elements described in any appendice may be applicable to various hardware, software, recording means, systems, and methods for recording software. [Explanation of Symbols]

[0090] 1. Video Processing Device 1a Video Processing Device 10. Video transmission device 11 Classification section 12 Generation part 13 Output section 14. Transmitter 20 Video playback device 21 Playback Department 30 Video Processing Systems 40 Video Processing Devices 41 processors 42 memory

Claims

1. A classification means for classifying each of the multiple image frames that make up a video into a distribution frame to be distributed and a non-distribution frame other than the distribution frame, A generation means that generates text describing the video content between the two distribution frames based on the non-distribution frame contained between the two distribution frames, A video processing device equipped with the following features.

2. The generation means generates text describing the video content between two adjacent distribution frames, based on the non-distribution frames that are included between two adjacent distribution frames in the playback order of the plurality of distribution frames. The video processing apparatus according to claim 1.

3. The generation means generates text as video content that describes the changes that occurred in the objects within the video between the two distribution frames. The video processing apparatus according to claim 1 or 2.

4. The generation means generates the text as information used to convey the video content to the user. The video processing apparatus according to claim 1 or 2.

5. The generation means generates subtitle information as text to be displayed together with the subsequent playback frame, which is in a later playback order among the two distribution frames. The video processing apparatus according to claim 1 or 2.

6. The classification means determines the distribution frame in accordance with the set frame rate, The generation means generates the text when the frame rate is less than or equal to a predetermined value. The video processing apparatus according to claim 1 or 2.

7. The generation means generates the text when the frame rate of the two distribution frames in a specific section becomes lower than or equal to a predetermined value compared to the frame rate of the plurality of image frames in that section. The video processing apparatus according to claim 1 or 2.

8. It has a video transmission device and a video playback device, The aforementioned video transmission device is A classification means for classifying each of the multiple image frames that make up a video into a distribution frame to be distributed and a non-distribution frame other than the distribution frame, A generation means that generates text describing the video content between the two distribution frames based on the non-distribution frame contained between the two distribution frames, The system includes a transmission means for transmitting the distribution frame and the text, The aforementioned video playback device is The system includes a playback means for playing back the distribution frame using the text received from the video transmission device. Video processing system.

9. Computers Each of the multiple image frames that make up the video is classified into a distribution frame to be distributed and a non-distribution frame other than the distribution frame. Based on the non-distributed frame contained between the two distribution frames, a text describing the video content between the two distribution frames is generated. Video processing methods.

10. A process that classifies each of the multiple image frames that make up a video into a distribution frame to be distributed and a non-distribution frame other than the distribution frame, A process to generate text describing the video content between the two distribution frames, based on the non-distribution frame contained between the two distribution frames, A program that causes a computer to execute something.

Citation Information

Patent Citations

  • Video conversion device, document conversion device, video conversion method, document conversion method, video conversion program and document conversion program

    JP2011243156A