Data augmentation device, data augmentation method, and program

The data augmentation device and method address the limitations of uniform speed changes in existing video data augmentation techniques by varying the frame extraction intervals, resulting in more realistic and varied video data that better represents real-world operator speeds, enhancing the accuracy and efficiency of analysis models.

WO2025126805A1PCT designated stage expired Publication Date: 2025-06-19NEC CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/041426
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-13
Filing Date
2024-11-22
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Existing data augmentation techniques for video data, such as extracting frames at equal intervals, result in a uniform change in the speed of the time flow, which may not accurately represent the varied working speeds of operators in real scenarios.

Method used

A data augmentation device and method that extract video frames from source video data at varying intervals, generating augmented video data where the separation between adjacent frames can differ, thereby creating uneven changes in the speed of the time flow, more accurately mimicking real-world variations in operator speed.

Benefits of technology

The proposed solution effectively increases the variability of video data used for training analysis models, allowing for more accurate representation and analysis of different working speeds, thus improving the model's training efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024041426_19062025_PF_FP_ABST
    Figure JP2024041426_19062025_PF_FP_ABST
Patent Text Reader

Abstract

A data augmentation device according to the present invention acquires source video data which is constituted by a plurality of video frames, extracts a plurality of video frames from the source video data, and generates augmented video data which is constituted by the plurality of extracted video frames. A first video frame and a second video frame which are adjacent to each other in the augmented video data are apart from each other by n frames in the source video data. A third video frame and a fourth video frame which are adjacent to each other in the augmented video data are apart from each other by m frames in the source video data. n and m differ from each other and are each an integer of 1 or more.
Need to check novelty before this filing date? Find Prior Art

Description

Data extension device, data extension method, and program

[0001] The present disclosure relates to a data extension device, a data extension method, and a program.

[0002] Systems have been developed that process data to generate new data, i.e., perform data augmentation. For example, Patent Literature 1 discloses a technology that extracts image data at equal intervals from video data and rearranges the extracted image data in reverse order to generate new video data.

[0003] Japanese Patent Application Laid-Open No. 2022-064460

[0004] In Patent Document 1, image data is extracted from the original video data at equal intervals. Therefore, the ratio of the speed of time flow between the new video data and the original video data is constant throughout the entire video data. The present disclosure has been made in consideration of this problem, and one of its purposes is to provide a new technology for data extension of video data.

[0005] A data extension device according to the present disclosure includes an acquisition means for acquiring source video data consisting of a plurality of video frames, and a generation means for extracting a plurality of video frames from the source video data and generating extended video data consisting of the extracted plurality of video frames. A first video frame and a second video frame that are adjacent to each other in the extended video data are separated by n frames in the source video data. A third video frame and a fourth video frame that are adjacent to each other in the extended video data are separated by m frames in the source video data, where n and m are integers that are different from each other and are equal to or greater than 1.

[0006] A data extension method according to the present disclosure is executed by a computer. The method includes an acquisition step of acquiring source video data consisting of a plurality of video frames, and a generation step of extracting a plurality of video frames from the source video data and generating extended video data consisting of the extracted plurality of video frames. A first video frame and a second video frame that are adjacent to each other in the extended video data are separated by n frames in the source video data. A third video frame and a fourth video frame that are adjacent to each other in the extended video data are separated by m frames in the source video data, where n and m are integers that are different from each other and are equal to or greater than 1.

[0007] The program of the present disclosure causes a computer to execute the data expansion method of the present disclosure.

[0008] According to the present disclosure, a new technique for performing data enhancement on video data is provided.

[0009] FIG. 1 is a diagram illustrating an overview of the operation of a data extension device. FIG. 2 is a block diagram illustrating an example of the functional configuration of a data extension device. FIG. 3 is a block diagram illustrating an example of the hardware configuration of a computer that realizes the data extension device. FIG. 4 is a flowchart illustrating an example of the flow of processing executed by the data extension device. FIG. 5 is a diagram illustrating an example of source video data composed of a sequence of multiple frames. FIG. 6 is a diagram illustrating an example of class information in a table format.

[0010] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In each drawing, the same or corresponding elements are designated by the same reference numerals, and duplicate explanations will be omitted as necessary for clarity. Furthermore, unless otherwise specified, predetermined values ​​such as predetermined values ​​and threshold values ​​are stored in advance in a storage device accessible from a device that uses the values. Furthermore, unless otherwise specified, the storage unit is composed of one or any number of storage devices.

[0011] <Overview> Fig. 1 is a diagram illustrating an example of an overview of the operation of the data expansion device 2000. Here, Fig. 1 is a diagram for facilitating understanding of the overview of the data expansion device 2000, and the operation of the data expansion device 2000 is not limited to that shown in Fig. 1.

[0012] The data extension device 2000 acquires source video data 10 and generates extended video data 30 from the source video data 10. Both the source video data 10 and the extended video data 30 are video data made up of a plurality of video frames.

[0013] The extended video data 30 is generated by extracting some video frames 12 from the source video data 10. When generating the extended video data 30 from the source video data 10, the order of the video frames 12 extracted from the source video data 10 is not changed. In other words, the order of the video frames 12 in the source video data 10 is the same as the order of the video frames 12 in the extended video data 30. The extended video data 30 can also be expressed as video data from which some video frames 12 have been deleted from the source video data 10.

[0014] The intervals at which video frames 12 are extracted from the source video data 10 (hereinafter referred to as extraction intervals) are not constant. Therefore, the extended video data 30 contains at least four video frames F1, F2, F3, and F4 that satisfy the following conditions:

[0015] (1) Video frames F1 and F2 are adjacent to each other in the extended video data 30, but are n frames apart in the source video data 10. (2) Video frames F3 and F4 are adjacent to each other in the extended video data 30, but are m frames apart in the source video data 10. (3) n and m are different integers greater than or equal to 1.

[0016] As will be described later, there are various methods for generating the extended video data 30 including the video frames F1 to F4. For example, the data extension device 2000 generates the extended video data 30 by extracting video frames 12 from the source video data 10 at random extraction intervals.

[0017] <Example of Effects> According to the data extension device 2000 of this embodiment, extended video data 30 is generated by extracting a plurality of video frames 12 from the source video data 10. Here, the extended video data 30 contains at least a pair of video frames 12 that are n frames apart in the source video data 10 and a pair of video frames 12 that are m frames (n ≠ m) apart in the source video data 10. By generating such extended video data 30, the speed of time flow in scenes captured in the extended video data 30 can be made non-uniformly faster compared to the speed of time flow in scenes captured in the source video data 10.

[0018] For example, assume that the frame rate of the source video data 10 is 30 frames per second (fps). Also, assume that a video frame 12 is extracted every two frames from one portion of the source video data 10, while a video frame 12 is extracted every five frames from another portion of the source video data 10. The portion of the extended video data 30 generated by extracting every two video frames 12 from the source video data 10 flows twice as fast in time as the source video data 10. On the other hand, the portion of the extended video data 30 generated by extracting every five video frames 12 from the source video data 10 flows five times as fast in time as the source video data 10.

[0019] For example, the data extension device 2000 is used to train an analytical model that analyzes video data capturing a scene in which work is being performed. Specifically, the data extension device 2000 performs data extension to generate extended video data 30 from source video data 10, thereby increasing the variety of video data used to train the analytical model. For example, the work may be various tasks such as tightening screws at a product manufacturing site.

[0020] The working speed may vary depending on the person (worker) performing the work. Therefore, training an analysis model requires a large amount of video data with different working speeds. However, preparing such a large amount of video data through actual filming is time-consuming and labor-intensive. Therefore, it is preferable to use the data extension device 2000 to generate extended video data 30 with different working speeds.

[0021] Here, by extracting video frames 12 at equal intervals from the source video data 10, it is possible to generate video data in which the working speed is uniformly faster than the working speed in the source video data 10. However, in reality, when comparing the working speeds of two workers, it is thought that the difference in working speed is often not uniform. For example, since some people are good at some tasks and others are not, the difference in working speed may not be uniform. Furthermore, it is thought that there are some tasks in which the working speed is likely to differ between workers, and other tasks in which the working speed is not likely to differ between workers.

[0022] Therefore, the data extension device 2000 does not make the interval at which video frames 12 are extracted from the source video data 10 constant throughout the source video data 10. By doing so, the speed of the work captured in the extended video data 30 can be made unevenly faster compared to the speed of the work captured in the source video data 10. Therefore, compared to when video frames 12 are extracted from the source video data 10 at uniform extraction intervals, it is possible to increase the variety of extended video data 30 that can be generated from the source video data 10.

[0023] The data expansion device 2000 of this embodiment will be described in more detail below.

[0024] 2 is a block diagram illustrating the functional configuration of a data expansion device 2000 according to embodiment 1. The data expansion device 2000 includes an acquisition unit 2020 and a generation unit 2040. The acquisition unit 2020 acquires source video data 10. The generation unit 2040 extracts a plurality of video frames 12 from the source video data 10 and generates extended video data 30 using the extracted plurality of video frames 12.

[0025] The extended video data 30 includes video frames F1, F2, F3, and F4. Video frames F1 and F2 are adjacent to each other in the extended video data 30, but are n frames apart in the source video data 10. Meanwhile, video frames F3 and F4 are adjacent to each other in the extended video data 30, but are m frames apart in the source video data 10, where n and m are different integers equal to or greater than 1.

[0026] <Example of Hardware Configuration> Each functional component of the data expansion device 2000 may be realized by hardware that realizes each functional component (e.g., a hardwired electronic circuit, etc.), or may be realized by a combination of hardware and software (e.g., a combination of an electronic circuit and a program that controls it, etc.). Below, a case where each functional component of the data expansion device 2000 is realized by a combination of hardware and software will be further described.

[0027] 3 is a block diagram illustrating an example of the hardware configuration of a computer 1000 that realizes the data expansion device 2000. The computer 1000 is any computer. For example, the computer 1000 is a stationary computer such as a PC (Personal Computer) or a server machine. Alternatively, the computer 1000 may be a portable computer such as a smartphone or a tablet terminal. The computer 1000 may be a dedicated computer designed to realize the data expansion device 2000, or may be a general-purpose computer.

[0028] For example, by installing a predetermined application on the computer 1000, the computer 1000 realizes each function of the data expansion device 2000. The application is configured with a program for realizing each functional component of the data expansion device 2000. The program can be acquired by any method. For example, the program can be acquired from a storage medium on which the program is stored. The storage medium on which the program is stored can be any storage medium such as a DVD (Digital Versatile Disk) or a USB (Universal Serial Bus) memory. Alternatively, the program can be acquired by downloading the program from a server device that manages the storage device on which the program is stored.

[0029] The computer 1000 has a bus 1020, a processor 1040, a memory 1060, a storage device 1080, an input / output interface 1100, and a network interface 1120. The bus 1020 is a data transmission path for the processor 1040, the memory 1060, the storage device 1080, the input / output interface 1100, and the network interface 1120 to transmit and receive data to and from each other. However, the method of connecting the processor 1040 and the like to each other is not limited to bus connection.

[0030] The processor 1040 is a processor such as a central processing unit (CPU), a graphics processing unit (GPU), or a field-programmable gate array (FPGA). The memory 1060 is a main storage device realized using a random access memory (RAM) or the like. The storage device 1080 is an auxiliary storage device realized using a hard disk, a solid state drive (SSD), a memory card, a read only memory (ROM), or the like.

[0031] The input / output interface 1100 is an interface for connecting the computer 1000 to an input / output device. For example, the input / output interface 1100 is connected to an input device such as a keyboard and an output device such as a display device.

[0032] The network interface 1120 is an interface for connecting the computer 1000 to a network. This network may be a LAN (Local Area Network) or a WAN (Wide Area Network).

[0033] The storage device 1080 stores a program (a program that realizes the above-mentioned application) that realizes each functional component of the data expansion device 2000. The processor 1040 reads this program into the memory 1060 and executes it to realize each functional component of the data expansion device 2000.

[0034] The data extension device 2000 may be realized by one computer 1000 or by multiple computers 1000. In the latter case, the configurations of the computers 1000 do not need to be the same, and can be different from each other.

[0035] 4 is a flowchart illustrating the flow of processing executed by the data extension device 2000 according to embodiment 1. The acquisition unit 2020 acquires the source video data 10 (S102). The generation unit 2040 extracts video frames 12 from the source video data 10 and generates extended video data 30 (S104).

[0036] <Acquisition of Source Video Data 10: S102> The acquisition unit 2020 acquires the source video data 10. Here, various methods can be used to acquire the video data to be processed. For example, the source video data 10 is assumed to be stored in advance in an arbitrary storage device in a format that allows it to be acquired by the data expansion device 2000. In this case, the acquisition unit 2020 acquires the source video data 10 by reading the source video data 10 from the storage device.

[0037] Alternatively, for example, the acquisition unit 2020 acquires the source video data 10 by receiving the source video data 10 transmitted from another device. The device that transmits the source video data 10 is, for example, the device that generated the source video data 10.

[0038] <Generation of Extended Video Data 30: S104> The generation unit 2040 extracts video frames 12 from the source video data 10 to generate extended video data 30 (S104). The generation unit 2040 extracts video frames 12 so that the extraction interval is not constant throughout the source video data 10. In other words, the video frames 12 are extracted so that the extraction interval is different in at least two locations, such as between video frames F1 and F2 and between video frames F3 and F4. Several specific methods for doing this are described below.

[0039] <<Specific Example 1 of Extraction Method>> For example, the generation unit 2040 randomly determines the extraction interval each time it extracts a video frame 12 from the source video data 10. The random determination of the extraction interval is performed, for example, using a predetermined probability distribution. That is, the generation unit 2040 extracts a value from the predetermined probability distribution and uses the extracted value as the extraction interval.

[0040] Various probability distributions can be adopted for determining the extraction interval. For example, the generation unit 2040 can use the uniform distribution defined by the following formula (1), the sigmoid defined by formula (2), the exponential distribution defined by formula (3), the Poisson distribution defined by formula (4), or the Gaussian distribution defined by formula (5). s represents a random variable. k is a parameter set in advance. a is a parameter set in advance.

[0041] The generation unit 2040 extracts a value s from the probability distribution and uses this value as the extraction interval. However, when the value s extracted from the probability distribution is not an integer, the generation unit 2040 converts s to an integer by any method such as applying a Gaussian function. Then, the generation unit 2040 uses the obtained integer as the extraction interval.

[0042] The generation unit 2040 may extract values from two or more probability distributions and use the statistical value of the extracted multiple values as the extraction interval. As the statistical value, for example, the simple average, weighted average, maximum value, or minimum value can be used.

[0043] For example, the generation unit 2040 uses the weighted average w1*Pg(s;k1,a)+w2*Pg(s;k2,a) of the values calculated from two Gaussian distributions Pg(s;k1,a) and Pg(s;k2,a) with different average values k as the extraction interval. Here, it is assumed that k1 < k2. Also, w1 and w2 are the weights given to the respective Gaussian distributions, and w1 + w2 = 1.

[0044] By using the above-mentioned weighted average, the extraction interval is highly likely to be a value between k1 and k2. Therefore, while changing the extraction interval, it is possible to make the extraction interval concentrate within or around the intended range.

[0045] <<Specific Example 2 of Extraction Method>> Assume here that the source video data 10 is composed of a sequence of frames each belonging to a different class. The sequence of frames is time-series data composed of a plurality of video frames that are consecutive in time series. Furthermore, all video frames 12 that belong to the same frame sequence belong to the same class.

[0046] 5 is a diagram illustrating source video data 10 that is composed of a sequence of multiple frames. In FIG. 5, the source video data 10 includes a sequence of multiple frames 20.

[0047] A frame sequence 20 is made up of a plurality of consecutive video frames 12 that belong to the same class. For example, the source video data 10 in Figure 5 includes, in this order, a frame sequence 20-1 made up of a plurality of video frames 12 that belong to class C1, a frame sequence 20-2 made up of a plurality of video frames 12 that belong to class C2, and a frame sequence 20-3 made up of a plurality of video frames 12 that belong to class C3.

[0048] Hereinafter, a frame sequence 20 consisting of a plurality of video frames 12 belonging to class C will also be referred to as a "frame sequence 20 belonging to class C." In this case, it will also be expressed as "the frame sequence 20 belongs to class C." Furthermore, the class to which the frame sequence 20 belongs will also be referred to as the "class of the frame sequence 20." If the frame sequence 20 belongs to class C, the class of the frame sequence 20 is C.

[0049] Here, the source video data 10 includes at least two frame sequences 20 that belong to different classes. However, the source video data 10 may also include two or more frame sequences 20 that belong to the same class. For example, in a case where the source video data 10 includes three frame sequences 20, namely, frame sequence 20-1 to frame sequence 20-3, frame sequence 20-1 and frame sequence 20-3 may belong to class C1, and frame sequence 20-2 may belong to class C2.

[0050] The class of the frame sequence 20 is determined, for example, according to the characteristics of the scene captured in the frame sequence 20. For example, assume that the source video data 10 contains a scene in which a worker is performing a task. In this case, for example, the class of the frame sequence 20 represents the type of task being performed by the worker in the scene captured in the frame sequence 20.

[0051] For example, suppose a video camera captures workers performing tasks of type A1, type A2, and type A3. The video data obtained by the capture is treated as source video data 10. In this case, A1, A2, and A3 are each treated as a class. The source video data 10 is divided into three frame sequences 20: a frame sequence 20 capturing the state of task A1, a frame sequence 20 capturing the state of task A2, and a frame sequence 20 capturing the state of task A3.

[0052] Assuming that the source video data 10 includes a plurality of frame sequences 20 in this way, the generation unit 2040, for example, determines an extraction interval for each frame sequence 20. There are various methods for determining an extraction interval for each frame sequence 20.

[0053] For example, the generation unit 2040 determines the extraction interval for each frame sequence 20 and keeps the extraction interval constant within the frame sequence 20. For example, assume that the source video data 10 is composed of three frame sequences V1, V2, and V3. In this case, the generation unit 2040 determines extraction intervals s1, s2, and s3 for the frame sequences V1, V2, and V3, respectively.

[0054] In this case, the generation unit 2040 extracts video frames 12 from the frame sequence V1 at a fixed extraction interval s1. That is, one video frame 12 is extracted from every s1 frames in the frame sequence V1. Similarly, the generation unit 2040 extracts video frames 12 from the frame sequence V2 at an extraction interval s2. Furthermore, the generation unit 2040 extracts video frames 12 from the frame sequence V3 at an extraction interval s3.

[0055] For example, the above-mentioned probability distribution can be used to determine the extraction interval for each frame sequence 20. In this case, a common probability distribution may be used for all frame sequences 20, or a different probability distribution may be used for each frame sequence 20. In the latter case, the generation unit 2040 determines the extraction interval for each frame sequence 20 using the probability distribution corresponding to that frame sequence 20.

[0056] There are various methods for determining the probability distribution corresponding to each frame sequence 20. For example, the probability distribution corresponding to each frame sequence 20 may be predetermined, may be determined randomly, or may be specified by the user of the data expansion device 2000.

[0057] The extraction interval may be determined for each class, rather than for each frame sequence 20. For example, assume that the source video data 10 is composed of frame sequences V1, V2, and V3, and that frame sequences V1 and V3 belong to class C1, and frame sequence V2 belongs to class C2. When the extraction interval is determined for each class, the generation unit 2040 determines extraction intervals s1 and s2 for each of classes C1 and C2. The generation unit 2040 extracts video frames 12 from frame sequence V1 at extraction interval s1, extracts video frames 12 from frame sequence V2 at extraction interval s2, and extracts video frames 12 from frame sequence V3 at extraction interval s1. In this way, video frames 12 are extracted at the same extraction interval s1 from frame sequences V1 and V3, which belong to the same class.

[0058] The sampling interval for each class can be determined using, for example, the above-mentioned probability distribution. In this case, a common probability distribution may be used for all classes, or a different probability distribution may be used for each class. In the latter case, the generation unit 2040 determines the sampling interval for each class using the probability distribution corresponding to that class.

[0059] There are various methods for determining the probability distribution corresponding to each class. For example, the probability distribution corresponding to each class may be predetermined, may be determined randomly, or may be specified by the user of the data expansion device 2000.

[0060] Determining the extraction interval for each class has the advantage that the differences between the classes can be reflected in the extended video data 30. For example, suppose that the classes represent types of work. In this case, it is considered that each worker has some tasks that they are good at and some tasks that they are not good at. Therefore, it is considered that differences in work speed are unlikely to occur between workers of the same type of work, but are likely to occur between workers of different types of work.

[0061] Therefore, by determining the task interval for each type of task, it is possible to generate extended video data 30 that simulates the characteristics of the worker, such as "making the task speed relatively fast for tasks that the worker is good at, and making the task speed relatively slow for tasks that the worker is not good at." Therefore, extended video data 30 having characteristics closer to those of real video data can be generated from the source video data 10. Furthermore, by generating extended video data 30 having characteristics closer to those of real video data in this way, when the extended video data 30 is used to train an analysis model, the accuracy of the analysis model can be improved.

[0062] The generation unit 2040 may vary the extraction interval within a single frame sequence. In this case, for example, the generation unit 2040 determines the extraction interval each time a video frame 12 is extracted using a different probability distribution for each frame sequence. The different probability distributions for each frame sequence may be of different types, or may be probability distributions of the same type but with different parameters.

[0063] For example, in the case where the source video data 10 is composed of a sequence of frames V1, V2, and V3 as described above, a Gaussian distribution Pg(s;k1,a1) is used for the sequence of frames V1, a Gaussian distribution Pg(s;k2,a2) is used for the sequence of frames V2, and a Gaussian distribution Pg(s;k3,a3) is used for the sequence of frames V3. When extracting multiple video frames 12 from the sequence of frames V1, the generation unit 2040 repeatedly extracts values ​​from the Gaussian distribution Pg(s;k1,a1) to determine the extraction interval each time a video frame 12 is extracted. For the sequence of frames V2, a similar process is performed using the Gaussian distribution Pg(s;k2,a2). For the sequence of frames V3, a similar process is performed using the Gaussian distribution Pg(s;k3,a3).

[0064] When varying the sampling interval within the frame sequence 20, different probability distributions may be used for each class, rather than for each frame sequence 20. For example, as described above, assume that the source video data 10 is composed of frame sequences V1, V2, and V3, and that the frame sequences V1 and V3 belong to class C1, and the frame sequence V2 belongs to class C2. In this case, for example, a Gaussian distribution Pg(s;k1,a1) is used to determine the sampling interval for class C1, and a Gaussian distribution Pg(s;k2,a2) is used to determine the sampling interval for class C2.

[0065] That is, when the generation unit 2040 extracts video frames 12 from each of the frame sequences V1 and V3 belonging to class C1, it extracts a value from the Gaussian distribution Pg(s;k1,a1) and determines the extraction interval each time a video frame 12 is extracted. On the other hand, when the generation unit 2040 extracts video frames 12 from the frame sequence V2 belonging to class C2, it extracts a value from the Gaussian distribution Pg(s;k2,a2) and determines the extraction interval each time a video frame 12 is extracted.

[0066] Here, in order to handle the frame sequence 20, it is necessary to be able to identify each frame sequence 20 from the source video data 10. Therefore, for example, the data extension device 2000 acquires information (hereinafter, referred to as class information) that can identify the position of each frame sequence 20 in the source video data 10.

[0067] For example, the class information indicates, for each video frame 12 included in the source video data 10, a correspondence between its identification information (e.g., frame number) and the identification information of the class to which the video frame 12 belongs. Alternatively, for each frame sequence 20 included in the source video data 10, the class information may indicate the identification information of either or both of the first and last video frames 12.

[0068] 6 is a diagram illustrating an example of class information in table format. Table 200 indicates, for each video frame 12, the class to which that video frame 12 belongs. More specifically, the table indicates, in association with the identification information of the video frame 12 (frame identification information 202), the identification information of the class to which that video frame 12 belongs (class identification information 204).

[0069] On the other hand, the table 300 indicates, for each frame sequence 20, the class to which that frame sequence 20 belongs. More specifically, for each frame sequence 20, the table 300 indicates the identification information of the class to which that frame sequence 20 belongs (class identification information 306) in association with a combination of the identification information of the first video frame 12 (first frame identification information 302) and the identification information of the last video frame 12 (last frame identification information 304).

[0070] The class information may be information that is integrated with the source video data 10, or may be information that is separate from the source video data 10. In the former case, for example, identification information of the class to which each video frame 12 included in the source video data 10 belongs is added as metadata. When the source video data 10 and the class information are configured separately, for example, the acquisition unit 2020 further acquires the class information for the source video data 10 in addition to the source video data 10. The method of acquiring the class information is the same as the method of acquiring the source video data 10.

[0071] <Output of Results> The data extension device 2000 outputs the execution results. Hereinafter, information output from the data extension device 2000 will be referred to as output information. The output information includes at least the extended video data 30. Furthermore, if the class information is configured separately from the video data, the output information further includes the class information of the extended video data 30. Here, if multiple pieces of extended video data 30 are generated, the output information includes multiple combinations of the extended video data 30 and the class information.

[0072] There are various methods for generating class information corresponding to the extended video data 30. For example, assume that the configuration of the class information is represented by a table 200. In this case, the generation unit 2040 extracts records of each video frame 12 extracted from the source video data 10 from the class information of the source video data 10. Then, the generation unit 2040 generates a table consisting of the extracted records as the class information of the extended video data 30.

[0073] On the other hand, it is assumed that the configuration of the class information is represented by a table 300. In this case, the generating unit 2040 generates the class information for the extended video data 30 by generating records of the table 300 for each frame sequence of the extended video data 30.

[0074] For example, suppose the source video data 10 includes a frame sequence V1, and a frame sequence of the extended video data 30 consisting of a plurality of video frames 12 extracted from the frame sequence V1 is called a frame sequence D1.

[0075] In this case, the generation unit 2040 generates a record in the table 300 for the frame sequence D1. The first frame identification information 302 of this record indicates the identification information of the first video frame 12 of the frame sequence D1 (i.e., the video frame 12 extracted first from the frame sequence V1). The last frame identification information 304 of this record indicates the identification information of the last video frame 12 of the frame sequence D1 (i.e., the video frame 12 extracted last from the frame sequence V1). The class identification information 306 of this record indicates the identification information of the class of the frame sequence D1 (i.e., the class of the frame sequence V1).

[0076] The generating unit 2040 performs the same process on each frame sequence included in the extended video data 30 to generate class information configured in a table 300 for the extended video data 30 .

[0077] The output information may be output in any manner. For example, the data extension device 2000 may store the output information in any storage device. Alternatively, the data extension device 2000 may transmit the output information to any device. For example, the destination device may be a device that uses the extended video data 30 to train a classifier that identifies the class of each video frame included in a frame sequence.

[0078] Although the present disclosure has been described above with reference to the embodiments, the present disclosure is not limited to the above-described embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure. Furthermore, each embodiment can be combined with other embodiments as appropriate.

[0079] Each drawing is merely an example for describing one or more embodiments. Each drawing may not relate to only one particular embodiment, but may also relate to one or more other embodiments. As will be understood by those skilled in the art, various features or steps described with reference to any one drawing can be combined with features or steps shown in one or more other drawings to create, for example, an embodiment not explicitly shown or described. Not all features or steps shown in any one drawing are necessary to describe an exemplary embodiment, and some features or steps may be omitted. The order of steps described in any drawing may be changed as appropriate.

[0080] In the present disclosure, a program includes a set of instructions (or software code) that, when loaded into a computer, causes the computer to perform one or more functions described in the embodiments. The program may be stored in a non-transitory computer-readable medium or a tangible storage medium. By way of example and not limitation, computer-readable medium or tangible storage medium includes random-access memory (RAM), read-only memory (ROM), flash memory, solid-state drive (SSD) or other memory technology, CD-ROM, digital versatile disc (DVD), Blu-ray disc or other optical disk storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage device. The program may also be transmitted on a transitory computer-readable medium or communication medium. By way of example and not limitation, transitory computer-readable medium or communication medium includes electrical, optical, acoustic, or other forms of propagated signals.

[0081] Some or all of the above embodiments can be described as, but are not limited to, the following supplementary notes. (Supplementary Note 1) A data extension device comprising: an acquisition means for acquiring source video data consisting of a plurality of video frames; and a generation means for extracting a plurality of video frames from the source video data and generating extended video data consisting of the extracted plurality of video frames, wherein a first video frame and a second video frame that are adjacent to each other in the extended video data are n frames apart in the source video data, and a third video frame and a fourth video frame that are adjacent to each other in the extended video data are m frames apart in the source video data, where n and m are integers greater than or equal to 1 and are not equal to each other. (Supplementary Note 2) The data extension device according to Supplementary Note 1, wherein the generation means extracts each of the plurality of video frames from the source video data at random intervals. (Supplementary Note 3) The data extension device according to Supplementary Note 2, wherein the generation means extracts a value from each of a plurality of probability distributions and uses a statistic of the extracted values ​​as the interval for extracting the video frames from the source video data. (Supplementary Note 4) The data extension device according to Supplementary Note 1, wherein the source video data includes a plurality of frame sequences each consisting of a plurality of consecutive video frames belonging to the same class, wherein adjacent frame sequences in the source video data belong to different classes, and wherein an interval at which the video frames are extracted from a first frame sequence of the source video data is different from an interval at which the video frames are extracted from a second frame sequence of the source video data. (Supplementary Note 5) The data extension device according to Supplementary Note 4, wherein the generation means determines an extraction interval for each of the frame sequences, and extracts the video frames from the frame sequences at the extraction interval determined for that frame sequence. (Supplementary Note 6) The data extension device according to Supplementary Note 4, wherein, each time the generation means extracts a video frame from the frame sequence, the generation means determines the interval at which the video frames are extracted from the frame sequence using a probability distribution corresponding to that frame sequence.(Supplementary Note 7) The data extension device according to Supplementary Note 4, wherein the generation means determines an extraction interval for each class, and extracts the video frames from the frame sequence at the extraction interval determined for the class to which the frame sequence belongs. (Supplementary Note 8) The data extension device according to Supplementary Note 4, wherein the generation means determines the extraction interval for extracting the video frames from the frame sequence, each time a video frame is extracted from the frame sequence, using a probability distribution corresponding to the class to which the frame belongs. (Supplementary Note 9) A computer-executed data extension method, comprising: an acquisition step of acquiring source video data consisting of a plurality of video frames; and a generation step of extracting a plurality of video frames from the source video data and generating extended video data consisting of the extracted plurality of video frames, wherein a first video frame and a second video frame that are adjacent to each other in the extended video data are spaced apart by n frames in the source video data, and a third video frame and a fourth video frame that are adjacent to each other in the extended video data are spaced apart by m frames in the source video data, wherein n and m are unequal integers of 1 or greater. (Supplementary Note 10) A program that causes a computer to execute an acquisition step of acquiring source video data consisting of a plurality of video frames; and a generation step of extracting a plurality of video frames from the source video data and generating extended video data consisting of the extracted plurality of video frames, wherein a first video frame and a second video frame that are adjacent to each other in the extended video data are n frames apart in the source video data, and a third video frame and a fourth video frame that are adjacent to each other in the extended video data are m frames apart in the source video data, and n and m are integers greater than or equal to 1 and are not equal to each other.

[0082] Some or all of the elements (e.g., configurations and functions) described in Supplementary Notes 2 to 8 that are dependent on Supplementary Note 1 may also be dependent on Supplementary Notes 9 and 10 in the same dependency relationship as Supplementary Notes 2 to 8. Some or all of the elements described in any Supplementary Note may be applied to various hardware, software, recording means for recording software, systems, and methods.

[0083] This application claims priority based on Japanese Patent Application No. 2023-210646, filed December 13, 2023, the disclosure of which is incorporated herein in its entirety by reference.

[0084] REFERENCE SIGNS LIST 10 Source video data 12 Video frame 20 Frame sequence 30 Extended video data 200 Table 202 Frame identification information 204 Class identification information 300 Table 302 First frame identification information 304 Last frame identification information 306 Class identification information 1000 Computer 1020 Bus 1040 Processor 1060 Memory 1080 Storage device 1100 Input / output interface 1120 Network interface 2000 Data extension device 2020 Acquisition unit 2040 Generation unit

Claims

1. A data extension device comprising: an acquisition means for acquiring source video data consisting of a plurality of video frames; and a generation means for extracting a plurality of video frames from the source video data and generating extended video data consisting of the extracted plurality of video frames, wherein a first video frame and a second video frame that are adjacent to each other in the extended video data are n frames apart in the source video data, and a third video frame and a fourth video frame that are adjacent to each other in the extended video data are m frames apart in the source video data, where n and m are different integers of 1 or greater.

2. The data extension device according to claim 1, wherein said generating means extracts a plurality of said video frames from said source video data at random intervals.

3. The data extension device according to claim 2, wherein said generating means extracts a value from each of a plurality of probability distributions and utilizes a statistic of said extracted values ​​as an interval for extracting said video frames from said source video data.

4. The data extension device of claim 1, wherein the source video data includes a plurality of frame sequences each consisting of a plurality of consecutive video frames belonging to the same class, wherein in the source video data, adjacent frame sequences belong to different classes, and wherein the interval at which the video frames are extracted from the first frame sequence of the source video data is different from the interval at which the video frames are extracted from the second frame sequence of the source video data.

5. A data extension device according to claim 4, wherein said generating means determines an extraction interval for each of said frame sequences, and extracts said video frames from said frame sequence at said extraction interval determined for said frame sequence.

6. The data extension device according to claim 4, wherein the generating means determines an interval for extracting the video frames from the frame sequence using a probability distribution corresponding to the frame sequence each time the generating means extracts the video frames from the frame sequence.

7. The data expansion device according to claim 4, wherein the generating means determines an extraction interval for each of the classes, and extracts the video frames from the frame sequence at the extraction interval determined for the class to which the frame sequence belongs.

8. A data extension device as described in claim 4, wherein the generating means determines an extraction interval for extracting the video frames from the frame sequence each time the video frames are extracted from the frame sequence using a probability distribution corresponding to the class to which the frame belongs.

9. A data extension method executed by a computer, comprising: an acquisition step of acquiring source video data consisting of a plurality of video frames; and a generation step of extracting a plurality of video frames from the source video data and generating extended video data consisting of the extracted plurality of video frames, wherein a first video frame and a second video frame that are adjacent to each other in the extended video data are n frames apart in the source video data, and a third video frame and a fourth video frame that are adjacent to each other in the extended video data are m frames apart in the source video data, where n and m are different integers greater than or equal to 1.

10. A program that causes a computer to execute an acquisition step of acquiring source video data consisting of a plurality of video frames; and a generation step of extracting a plurality of video frames from the source video data and generating extended video data consisting of the extracted plurality of video frames, wherein a first video frame and a second video frame that are adjacent to each other in the extended video data are n frames apart in the source video data, and a third video frame and a fourth video frame that are adjacent to each other in the extended video data are m frames apart in the source video data, where n and m are different integers of 1 or greater.

Citation Information

Patent Citations

  • Image retrieval model training method and device, computing equipment and storage medium

    CN117216305A

  • Device and method for processing video and computer program

    JP2006270405A

  • Moving image processing system, moving image processing method, and program

    JP2012151705A

  • Recording device

    JP2018195946A

  • Learning model generation device, learning model, behavior recognition device, learning data generation device, and learning data generation method

    JP2022064460A