Guidance data generation device, guidance data generation method, and program

The training data generation device addresses the challenge of manually recording task instructions by automatically detecting and generating teaching data from image and audio sequences, enhancing efficiency and usability for task analysis and machine learning.

WO2026014265A1PCT designated stage Publication Date: 2026-01-15NEC CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/023177
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-09
Filing Date
2025-06-27
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

Existing technologies fail to effectively analyze and automatically generate training data from image sequences capturing task instructions, requiring manual effort and time for recording past training information.

Method used

A training data generation device and method that detects and generates teaching data from image frames containing task instructions, utilizing machine learning models to identify and group instruction frames, and optionally incorporates audio data for enhanced data generation.

Benefits of technology

Automatically generates training data with reduced time and effort, providing valuable information for task understanding and machine learning model training, while offering audio and textual representations of the training content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025023177_15012026_PF_FP_ABST
    Figure JP2025023177_15012026_PF_FP_ABST
Patent Text Reader

Abstract

This guidance data generation device: acquires a sequence of image frames capturing work; detects, from among the sequence of image frames, a guidance image frame, which is an image frame capturing guidance relating to the work; and generates guidance data relating to the guidance captured in the sequence of image frames on the basis of the guidance image frame.
Need to check novelty before this filing date? Find Prior Art

Description

Training data generating device, training data generating method, and program

[0001] The present disclosure relates to a training data generating device, a training data generating method, and a program.

[0002] A technology has been developed to identify the type of work captured in an image sequence. Patent Document 1 discloses a technology for detecting abnormal work by identifying the work being performed by a worker captured in a work video.

[0003] International Publication No. 2023 / 058164

[0004] Patent Document 1 does not assume that instruction regarding a task may be included in the task video. The present disclosure has been made in consideration of this problem, and one of its purposes is to provide a new technique for analyzing the task captured in the image sequence.

[0005] The teaching data generation device according to the present disclosure includes an acquisition means for acquiring a sequence of image frames in which a task is captured, a detection means for detecting teaching image frames from the sequence of image frames, which are image frames in which instruction regarding the task is captured, and a generation means for generating teaching data regarding the instruction captured in the sequence of image frames based on the teaching image frames.

[0006] The teaching data generation method of the present disclosure includes an acquisition step of acquiring a sequence of image frames in which a task is captured, a detection step of detecting teaching image frames from the sequence of image frames, which are image frames in which instruction regarding the task is captured, and a generation step of generating teaching data regarding the instruction captured in the sequence of image frames based on the teaching image frames.

[0007] The program of the present disclosure causes a computer to execute an acquisition step of acquiring a sequence of image frames in which a task is being imaged, a detection step of detecting instruction image frames from the sequence of image frames, which are image frames in which instruction regarding the task is being imaged, and a generation step of generating instruction data regarding the instruction imaged in the sequence of image frames based on the instruction image frames.

[0008] The present disclosure provides a new technique for performing analysis on an activity captured in an image sequence.

[0009] FIG. 1 is a diagram illustrating an example of an outline of the operation of a training data generating device. FIG. 2 is a block diagram illustrating an example of the functional configuration of a training data generating device. FIG. 3 is a block diagram illustrating an example of the hardware configuration of a computer that realizes the training data generating device. FIG. 4 is a flowchart illustrating an example of the flow of processing executed by the training data generating device. FIG. 5 is a second diagram illustrating an example of an outline of the operation of a training data generating device. FIG. 6 is a second block diagram illustrating an example of the functional configuration of a training data generating device. FIG. 7 is a second flowchart illustrating an example of the flow of processing executed by the training data generating device.

[0010] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In each drawing, identical or corresponding elements are designated by the same reference numerals, and duplicate explanations will be omitted as necessary for clarity. Furthermore, unless otherwise specified, predetermined values ​​such as predetermined values ​​and thresholds are pre-stored in a storage device accessible from a device that uses the values. Furthermore, unless otherwise specified, the storage unit is composed of one or more arbitrary number of storage devices. Furthermore, unless otherwise specified, various models such as neural networks and support vector machines can be used as machine learning models.

[0011] [First Embodiment] <Overview> Fig. 1 is a diagram illustrating an example of an outline of the operation of a training data generation device 2000. Here, Fig. 1 is a diagram for facilitating understanding of the outline of the training data generation device 2000, and the operation of the training data generation device 2000 is not limited to that shown in Fig. 1.

[0012] The training data generation device 2000 analyzes an image frame sequence 10. The image frame sequence 10 is made up of a plurality of image frames 12 arranged in time series. In other words, the image frame sequence 10 is a frame sequence in which a plurality of image frames 12 are arranged in time series order (in ascending order of frame numbers).

[0013] Each image frame 12 belongs to one of a plurality of classes. Hereinafter, the class to which an image frame 12 belongs is also referred to as the "class of the image frame 12."

[0014] A scene in which a worker (a person performing a task) is being performed is captured in the image frame sequence 10. For example, the image frame sequence 10 is generated by capturing the worker's task with a video camera.

[0015] The class of an image frame 12 represents the type of work being performed by a worker in the scene captured in the image frame 12. For example, suppose a video camera captures a scene in which a worker is performing a type A1 work, a type A2 work, and a type A3 work. The video data obtained by the capture is then treated as an image frame sequence 10. In this case, A1, A2, and A3 are each treated as a class.

[0016] At this time, the worker may be given guidance regarding the work. For example, if the worker makes a mistake in the work, the instructor may provide guidance.

[0017] The training data generation device 2000 detects one or more image frames 12 in which instruction on a task is captured from the image frame sequence 10. Hereinafter, the image frames 12 in which instruction on a task is captured are referred to as training image frames 30. The training data generation device 2000 generates training data 40 relating to the content of the instruction based on the training image frames 30.

[0018] <Example of Effects> According to the teaching data generation device 2000 of this embodiment, teaching image frames 30 in which the state of teaching is captured are detected from the image frame sequence 10, and teaching data 40 relating to the content of the teaching is generated based on the teaching image frames 30. In this way, the teaching data generation device 2000 provides a new technology for automatically generating information relating to teaching from the image frame sequence in a technology for analyzing the work captured in the image frame sequence.

[0019] There are various advantages to automatically generating the training data 40 from the image frame sequence 10. For example, for a worker who is unfamiliar with a task, referring to the content of past training given to himself or other workers is useful for understanding the tricks of the task. However, manually recording information about past training requires time and effort. In this regard, the training data generation device 2000 automatically generates the training data 40, so that information about past training can be saved with little time and effort.

[0020] Furthermore, the training data generation device 2000 can acquire training image frames 30 depicting the state of work from the image frame sequence 10. As will be described later, it can also acquire a training image frame sequence consisting of a plurality of training image frames 30 in time series. The training image frames 30 and the training image frame sequence can also be used, for example, for training a machine learning model.

[0021] The training data generating device 2000 of this embodiment will be described in more detail below.

[0022] 2 is a block diagram illustrating an example of the functional configuration of the training data generation device 2000. The training data generation device 2000 has an acquisition unit 2020, a detection unit 2040, and a generation unit 2060. The acquisition unit 2020 acquires an image frame sequence 10. The detection unit 2040 detects a training image frame 30 from the image frame sequence 10. The generation unit 2060 generates training data 40 based on the training image frame 30.

[0023] <Example of Hardware Configuration> Each functional component of the training data generation device 2000 may be realized by hardware that realizes each functional component (e.g., a hardwired electronic circuit, etc.), or may be realized by a combination of hardware and software (e.g., a combination of an electronic circuit and a program that controls it, etc.). Below, a case where each functional component of the training data generation device 2000 is realized by a combination of hardware and software will be further described.

[0024] 3 is a block diagram illustrating an example of the hardware configuration of a computer 1000 that realizes the training data generation device 2000. The computer 1000 is any computer. For example, the computer 1000 is a stationary computer such as a PC (Personal Computer) or a server machine. Alternatively, the computer 1000 may be a portable computer such as a smartphone or a tablet terminal. The computer 1000 may be a dedicated computer designed to realize the training data generation device 2000, or may be a general-purpose computer.

[0025] For example, by installing a predetermined application on the computer 1000, the computer 1000 realizes each function of the training data generation device 2000. The application is configured with a program for realizing each functional component of the training data generation device 2000. The method for acquiring the program is arbitrary. For example, the program can be acquired from a storage medium on which the program is stored. The storage medium on which the program is stored may be any storage medium such as a DVD (Digital Versatile Disk) or a USB (Universal Serial Bus) memory. Alternatively, the program can be acquired by downloading the program from a server device that manages the storage device on which the program is stored.

[0026] The computer 1000 has a bus 1020, a processor 1040, a memory 1060, a storage device 1080, an input / output interface 1100, and a network interface 1120. The bus 1020 is a data transmission path for the processor 1040, the memory 1060, the storage device 1080, the input / output interface 1100, and the network interface 1120 to transmit and receive data to and from each other. However, the method of connecting the processor 1040 and the like to each other is not limited to bus connection.

[0027] The processor 1040 is a processor such as a central processing unit (CPU), a graphics processing unit (GPU), or a field-programmable gate array (FPGA). The memory 1060 is a main storage device realized using a random access memory (RAM) or the like. The storage device 1080 is an auxiliary storage device realized using a hard disk, a solid state drive (SSD), a memory card, a read only memory (ROM), or the like.

[0028] The input / output interface 1100 is an interface for connecting the computer 1000 to an input / output device. For example, the input / output interface 1100 is connected to an input device such as a keyboard and an output device such as a display device.

[0029] The network interface 1120 is an interface for connecting the computer 1000 to a network. This network may be a LAN (Local Area Network) or a WAN (Wide Area Network).

[0030] The storage device 1080 stores a program (a program that realizes the above-mentioned application) that realizes each functional component of the training data generation device 2000. The processor 1040 reads this program into the memory 1060 and executes it to realize each functional component of the training data generation device 2000.

[0031] The training data generation device 2000 may be realized by one computer 1000 or by multiple computers 1000. In the latter case, the configurations of the computers 1000 do not need to be the same, and can be different from each other.

[0032] 4 is a flowchart illustrating the flow of processing executed by the training data generation device 2000. The acquisition unit 2020 acquires an image frame sequence 10 (S102). The detection unit 2040 detects a training image frame 30 from the image frame sequence 10 (S104). The generation unit 2060 generates training data 40 based on the training image frame 30 (S106).

[0033] <Acquisition of Image Frame Sequence 10: S102> The acquisition unit 2020 acquires the image frame sequence 10. Here, various methods can be used to acquire the image frame sequence to be processed. For example, the image frame sequence 10 is stored in advance in an arbitrary storage device in a format that allows it to be acquired from the training data generation device 2000. In this case, the acquisition unit 2020 acquires the image frame sequence 10 by reading the image frame sequence 10 from the storage device.

[0034] Alternatively, for example, the acquisition unit 2020 may acquire the image frame sequence 10 by receiving the image frame sequence 10 transmitted from another device. The device that transmits the image frame sequence 10 is, for example, the device that generated the image frame sequence 10. If the image frame sequence 10 is video data, for example, the acquisition unit 2020 acquires the image frame sequence 10 from the video camera that generated the image frame sequence 10.

[0035] <Detection of Teaching Image Frames 30: S104> The detection unit 2040 detects teaching image frames 30 from the image frame sequence 10 (S104). Below, several examples of methods for detecting teaching image frames 30 will be described.

[0036] <<First Detection Method>> For example, the detection unit 2040 performs a process of identifying the type of task (i.e., class) captured in each image frame 12 included in the image frame sequence 10. The process of identifying the class is performed using, for example, a trained machine learning model. Hereinafter, this machine learning model will be referred to as a class identification model.

[0037] For example, the class identification model is configured to output a vector (hereinafter referred to as a class vector) representing the probability that an input image belongs to each of a plurality of predetermined classes. The class vector has the same number of elements as the number of predetermined classes. The i-th element of the class vector represents the probability that the input image belongs to the i-th class. The class of the image frame 12 is the class corresponding to the element with the largest value in the class vector obtained by inputting the image frame 12 into the class identification model.

[0038] When guidance is given during a task, the state of the task included in the image frame sequence 10 will be different from the state of the task when no guidance is given, and as a result, the probability that the class vector indicates the type of task becomes relatively small.

[0039] For example, suppose that image frame f1 captures a task of type C1, and image frame f2 captures instruction during the task of type C1. Furthermore, suppose that class vectors V1 and V2 are obtained by inputting image frames f1 and f2 into a class classification model.

[0040] In this case, both class vectors V1 and V2 have the largest value in the element corresponding to class C1, but the value of the element corresponding to class C1 in V2 is smaller than the value of the element corresponding to class C1 in V1.

[0041] Therefore, for example, a threshold Th1 is determined in advance as a threshold for distinguishing between cases where instruction is not being given and cases where instruction is being given. The detection unit 2040 obtains a class vector by inputting the image frame 12 into a class identification model. The detection unit 2040 identifies the class corresponding to the maximum element in the class vector as the class of the image frame 12.

[0042] Furthermore, the detection unit 2040 determines whether the value of the maximum element in the class vector is equal to or greater than a threshold value Th1. If the value of the maximum element in the class vector is equal to or greater than the threshold value Th1, the detection unit 2040 determines that the image frame 12 is not a teaching image frame 30. On the other hand, if the value of the maximum element indicated in the class vector is less than the threshold value Th1, the detection unit 2040 determines that the image frame 12 is a teaching image frame 30. If the identified class is Ck, the detected teaching image frame 30 is a teaching image frame 30 in which instruction on a task of class Ck is captured.

[0043] Here, when the value of the maximum element in the class vector is less than the threshold value Th1, the detection unit 2040 may further analyze the content of the image frame 12 to determine whether or not the image frame 12 is a teaching image frame 30. Specifically, the detection unit 2040 determines whether or not the image frame 12, for which the value of the maximum element in the class vector is determined to be less than the threshold value Th1, has the characteristics of a teaching image frame 30.

[0044] If the image frame 12 has the characteristics of a teaching image frame 30, the detection unit 2040 determines that the image frame 12 is a teaching image frame 30. On the other hand, if the image frame 12 does not have the characteristics of a teaching image frame 30, the detection unit 2040 determines that the image frame 12 is not a teaching image frame 30.

[0045] Various features can be adopted as the features of the instruction image frame 30. For example, features such as "two or more people are captured" or "a specific action is being performed" can be adopted.

[0046] The detection unit 2040 may further include a machine learning model (hereinafter, a discrimination model) that determines whether or not an image has characteristics of the teaching image frame 30. The discrimination model outputs, for example, a flag (hereinafter, a discrimination flag) that indicates whether or not an input image has characteristics of the teaching image frame 30.

[0047] The detection unit 2040 inputs the image frame 12, for which the value of the largest element in the class vector is determined to be less than the threshold Th1, into the discrimination model. If the discrimination flag indicates that the image frame 12 has the characteristics of a teaching image frame 30, the detection unit 2040 determines that the image frame 12 is a teaching image frame 30. On the other hand, if the discrimination flag indicates that the image frame 12 does not have the characteristics of a teaching image frame 30, the detection unit 2040 determines that the image frame 12 is not a teaching image frame 30.

[0048] <<<About Model Training>>> The class identification model is trained in advance so that it outputs a class vector in response to an input image. The training data used for this training consists of a combination of training images of tasks and ground truth class vectors. The ground truth class labels are, for example, one-hot vectors that indicate 1 for elements corresponding to the type of task captured in the training images and 0 for other elements.

[0049] A device for training a model (hereinafter referred to as a training device) obtains class vectors by inputting training images into a class discrimination model. The training device then calculates a loss based on the class vectors and ground truth class vectors, and updates the parameters of the class discrimination model based on the loss. The class discrimination model is trained by repeatedly updating the parameters of the class discrimination model using multiple training data.

[0050] The discrimination model is trained in advance to output a discrimination flag in response to an input image. The training data used for this training is composed of a combination of training images and ground truth discrimination flags. If the training images include images of instruction, the ground truth discrimination flag indicates that the image "has the characteristics of an instruction image frame 30" (for example, indicates 1). On the other hand, if the training images do not include images of instruction, the ground truth discrimination flag indicates that the image "does not have the characteristics of an instruction image frame 30" (for example, indicates 0).

[0051] The training device inputs training images into a discriminant model to obtain a discriminant flag. Furthermore, the training device calculates a loss based on the discriminant flag and a ground truth discriminant flag, and updates the parameters of the discriminant model based on the loss. The discriminant model is trained by repeatedly updating the parameters of the discriminant model using multiple training data.

[0052] <<Second Detection Method>> The class discrimination model described above may be configured to treat a "tutoring class" as one of the classes. If the class discrimination model treats N types of tasks, the class "tutoring" is treated as the (N+1)th class. In response to an input image of a teaching situation, the class discrimination model outputs a class vector with the largest element corresponding to the teaching class.

[0053] The detection unit 2040 obtains a class vector by inputting the image frame 12 into a class identification model. If the class corresponding to the largest element in the class vector is the training class, the detection unit 2040 determines that the image frame 12 is the training image frame 30. On the other hand, if the class corresponding to the largest element in the class vector is not the training class, the detection unit 2040 determines that the image frame 12 is not the training image frame 30.

[0054] <<Third Detection Method>> It is considered that instruction continues for a certain period of time, such as a few seconds or a few minutes, etc. Therefore, a sequence of instruction image frames (plurality of instruction image frames 30 consecutive in time series) showing the state of instruction can be detected from the image frame sequence 10.

[0055] Therefore, the detection unit 2040 may detect one or more teaching image frame sequences from the image frame sequence 10. For example, the detection unit 2040 detects one or more teaching image frames 30 from the image frame sequence 10 using the first detection method or the second detection method described above. The detection unit 2040 then groups together multiple teaching image frames 30 that are consecutive in time series and treats them as a single teaching image frame sequence. Note that when the class of each teaching image frame 30 is identified, it is preferable that the detection unit 2040 groups together multiple teaching image frames 30 that are consecutive in time series and belong to the same class and treats them as a single teaching image frame sequence.

[0056] Here, suppose that between two instruction image frame sequences there are a small number of image frames 12 that are determined not to be instruction image frames 30. In this case, there is a high probability that these small number of image frames 12 are actually instruction image frames 30 and, together with the preceding and following instruction image frame sequences, represent the state of instruction.

[0057] Therefore, the detection unit 2040 determines whether the number of image frames 12 existing between two instruction image frame sequences is equal to or less than a predetermined threshold. If the number of image frames 12 existing between the two instruction image frame sequences is equal to or less than the threshold, the detection unit 2040 treats the two instruction image frame sequences and all image frames 12 located between them as a single instruction image frame sequence. In this case, the detection unit 2040 changes each image frame 12 included in the instruction image frame sequence that was determined not to be an instruction image frame 30 to an instruction image frame 30.

[0058] Similarly, suppose that there are a threshold or less number of image frames 12 between a teaching image frame 30 and a teaching image frame sequence that are determined not to be teaching image frames 30. In this case, the detection unit 2040 treats the teaching image frame 30 and the teaching image frame sequence as a single teaching image frame sequence. In this case, the detection unit 2040 also changes each image frame 12 included in the teaching image frame sequence that was determined not to be a teaching image frame 30 to a teaching image frame 30. Furthermore, the class of these image frames 12 is changed to the class to which the teaching image frame sequence belongs.

[0059] The teaching image frame sequences that are combined into one using the above method may be limited to teaching image frame sequences that belong to the same class. Suppose two teaching image frame sequences belong to the same class and the number of image frames 12 between the two teaching image frame sequences is less than or equal to a threshold. In this case, the detection unit 2040 combines the two teaching image frame sequences and all image frames 12 located between them and treats them as a single teaching image frame sequence. At this time, the class of each image frame 12 located between the two teaching image frame sequences is changed to the class to which the teaching image frame sequence belongs. If two teaching image frame sequences belong to different classes, the two teaching image frame sequences will not be combined into one even if the number of image frames 12 between the two teaching image frame sequences is less than or equal to a threshold.

[0060] <Generation of Training Data 40: S106> The generating unit 2060 generates training data 40 based on the detected training image frames 30 (S106). For example, the generating unit 2060 generates training data 40 including one or more detected training image frames 30. The generating unit 2060 may further include additional information related to the training image frames 30 included in the training data 40 in the training data 40.

[0061] The additional information related to the teaching image frame 30 indicates, for example, the type of work represented by the teaching image frame 30. The type of work represented by the teaching image frame 30 is the class corresponding to the element with the largest value in the class vector output by the class identification information.

[0062] When a teaching image frame sequence is detected, the teaching data 40 may indicate the teaching image frame sequence. In this case, for example, the teaching data 40 may indicate, as additional information of the teaching image frame sequence, the class of the teaching image frame sequence or the period of teaching represented by the teaching image frame sequence. The start point of the teaching period indicates the generation time of the teaching image frame 30 located at the beginning of the teaching image frame sequence. On the other hand, the end point of the teaching period indicates the generation time of the teaching image frame 30 located at the end of the teaching image frame sequence.

[0063] The training data 40 may be output in any manner. For example, the training data generation device 2000 may store the training data 40 in any storage device. Alternatively, for example, the training data generation device 2000 may transmit the training data 40 to any device. Alternatively, for example, the training data generation device 2000 may display the training data 40 on any display device.

[0064] [Embodiment 2] Fig. 5 is a second diagram illustrating an outline of the operation of the training data generation device 2000. Here, Fig. 5 is a diagram for facilitating understanding of the outline of the training data generation device 2000, and the operation of the training data generation device 2000 is not limited to that shown in Fig. 5.

[0065] The training data generation device 2000 of the second embodiment acquires video data 50. The video data 50 includes a combination of an image frame sequence 10 and audio data 60. The image frame sequence 10 visually records the state of work. Meanwhile, the audio data 60 audibly records the state of work. For example, the video data 50 is generated by a video camera that is set up to record the state of work.

[0066] The training data generation device 2000 detects one or more training image frame sequences 80 from the image frame sequence 10. The training image frame sequence 80 is made up of a plurality of chronologically consecutive training image frames 30. Furthermore, the training data generation device 2000 extracts audio data (hereinafter referred to as training audio data 70) corresponding to the training image frame sequence 80 from the audio data 60.

[0067] For example, assume that the instruction image frame sequence 80 is composed of instruction image frames 30 from time t1 to time t2. In this case, the instruction audio data 70 corresponding to the instruction image frame sequence 80 is the portion of the audio data 60 from time t1 to time t2.

[0068] The training data generating device 2000 generates training data 40 based on the training voice data 70. For example, the training data 40 includes the training voice data 70 itself and text data (hereinafter referred to as training text) representing the content of the utterances included in the training voice data 70.

[0069] The instruction voice data 70 is generated based on the instruction image frame sequence 80. Therefore, generating the instruction data 40 based on the instruction voice data 70 is included in the scope of generating the instruction data 40 based on the instruction image frames 30.

[0070] <Example of Function and Effect> According to the teaching data generation device 2000 of this embodiment, a teaching image frame sequence 80 in which a teaching situation is captured is detected from the image frame sequence 10. Furthermore, teaching audio data 70 corresponding to the teaching image frame sequence 80 is detected from the audio data 60. Then, the teaching data 40 is generated based on the teaching audio data 70. Thus, according to the teaching data generation device 2000, the teaching data 40 is generated based on audio data that represents the content of the teaching. Therefore, according to the teaching data generation device 2000, information such as words and sentences that represent the content of the teaching can be obtained as the teaching data 40.

[0071] For example, assume that the training data 40 includes the training audio data 70. In this case, the user of the training data generation device 2000 can obtain audio representing words and sentences that represent the content of the training from the video data 50. Therefore, the user can easily understand the content of the training aurally.

[0072] Alternatively, for example, the training data 40 may include training text. In this case, the user of the training data generation device 2000 can obtain words and sentences expressing the content of the training as text from the video data 50. This allows the user to refer to the content of the training in text. Furthermore, text data has a smaller data size than image data and audio data. Therefore, when training text is used instead of the training image frames 30 and the training audio data 70, the data size of the training data 40 can be reduced.

[0073] <Example of Functional Configuration> Fig. 6 is a second block diagram illustrating the functional configuration of the training data generation device 2000. The functional configuration of the training data generation device 2000 in Fig. 6 is the same as the functional configuration of the training data generation device 2000 in Fig. 2 except that it includes an extraction unit 2080.

[0074] The detection unit 2040 of the second embodiment detects a plurality of teaching image frames 30 from the image frame sequence 10, thereby detecting one or more teaching image frame sequences 80 from the image frame sequence 10. The extraction unit 2080 extracts, for each teaching image frame sequence 80, teaching audio data 70 corresponding to that teaching image frame sequence 80 from the audio data 60. The generation unit 2060 generates teaching data 40 based on the teaching audio data 70.

[0075] <Example of Hardware Configuration> The hardware configuration of the training data generation device 2000 of the second embodiment is shown in, for example, Fig. 3 , similar to the hardware configuration of the training data generation device 2000 of the first embodiment. However, the storage device 1080 of the second embodiment stores a program for realizing the functions of the training data generation device 2000 of the embodiment.

[0076] <Processing Flow> Fig. 7 is a second flowchart illustrating the processing flow executed by the training data generation device 2000. The acquisition unit 2020 acquires the image frame sequence 10 and the audio data 60 (S202). The detection unit 2040 detects the training image frame sequence 80 from the image frame sequence 10 (S204). The extraction unit 2080 extracts audio data corresponding to the training image frame sequence 80 from the audio data 60 as training audio data 70 (S206). The generation unit 2060 generates training data 40 based on the training audio data 70 (S208).

[0077] <Acquisition of Image Frame Sequence 10 and Audio Data 60: S202> The acquisition unit 2020 acquires the image frame sequence 10 and the audio data 60 (S202). For example, the acquisition unit 2020 acquires the video data 50 including the image frame sequence 10 and the audio data 60 by a method similar to the various methods for acquiring the image frame sequence 10 described above. In this way, the image frame sequence 10 and the audio data 60 are acquired.

[0078] The acquisition unit 2020 may separately acquire the image frame sequence 10 and the audio data 60. In this case, the audio data 60 is acquired by a method similar to the various methods for acquiring the image frame sequence 10 described above.

[0079] <Detection of Teaching Image Frame Sequence 80: S204> The detection unit 2040 detects the teaching image frame sequence 80 from the image frame sequence 10 (S204). The method for detecting the teaching image frame sequence from the image frame sequence 10 is the same as that described in the first embodiment.

[0080] <Extraction of Guidance Audio Data 70: S206> For each of the instruction image frame sequences 80, the extraction unit 2080 extracts the instruction audio data 70 corresponding to the instruction image frame sequence 80 from the audio data 60 (S206). For example, the extraction unit 2080 extracts audio data from the audio data 60 covering the period from the start to the end of the instruction image frame sequence 80, and treats the extracted audio data as the instruction audio data 70.

[0081] <Generation of Training Data 40: S208> The generating unit 2060 generates the training data 40 based on the training audio data 70 (S208). For example, the generating unit 2060 generates the training data 40 including one or more pieces of training audio data 70. The training data 40 may further include additional information about the training audio data 70. The additional information about the training audio data 70 indicates, for example, the duration of the training audio data 70 (i.e., the period from the start to the end of the training audio data 70).

[0082] The instruction data 40 may include text data (i.e., instruction text) representing the content of the utterances included in the instruction audio data 70, together with the instruction audio data 70, or instead of the instruction audio data 70. In this case, the generation unit 2060 generates the instruction text by analyzing the instruction audio data 70.

[0083] To generate the training text, for example, a machine learning model (hereinafter referred to as a voice recognition model) that has been trained to generate text data representing the content of utterances included in the voice data is used. In this case, the generation unit 2060 inputs the training voice data 70 into the voice recognition model, and obtains the training text from the voice recognition model.

[0084] The training data 40 may further include a training image frame sequence 80 associated with the training audio data 70, the training text, or both. Extracting a combination of the training audio data 70 and the training image frame sequence 80 corresponds to extracting video data for a period representing the training from the video data 50.

[0085] The output mode of the training data 40 in the second embodiment is the same as the output mode of the training data 40 in the first embodiment. However, when the training data 40 includes the training audio data 70, the training audio data 70 can be output from a device that outputs audio, such as a speaker or headphones.

[0086] Although the present disclosure has been described above with reference to the embodiments, the present disclosure is not limited to the above-described embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure. Furthermore, each embodiment can be combined with other embodiments as appropriate.

[0087] Each drawing is merely an example for describing one or more embodiments. Each drawing may not relate to only one particular embodiment, but may also relate to one or more other embodiments. As will be understood by those skilled in the art, various features or steps described with reference to any one drawing can be combined with features or steps shown in one or more other drawings to create, for example, an embodiment not explicitly shown or described. Not all features or steps shown in any one drawing are necessary to describe an exemplary embodiment, and some features or steps may be omitted. The order of steps described in any drawing may be changed as appropriate.

[0088] In the present disclosure, a program includes instructions (or software code) that, when loaded into a computer, cause the computer to perform one or more functions described in the embodiments. The program may be stored in a non-transitory computer-readable medium or a tangible storage medium. By way of example and not limitation, computer-readable media or tangible storage media include random-access memory (RAM), read-only memory (ROM), flash memory, solid-state drive (SSD) or other memory technologies, CD-ROM, digital versatile disc (DVD), Blu-ray disc or other optical disk storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage device. The program may also be transmitted on a transitory computer-readable medium or communication medium. By way of example and not limitation, transitory computer-readable media or communication media include electrical, optical, acoustic, or other forms of propagated signals.

[0089] Some or all of the above embodiments can be described as, but are not limited to, the following supplementary notes. (Supplementary Note 1) A training data generation device comprising: acquisition means for acquiring an image frame sequence in which a task is captured; detection means for detecting, from the image frame sequence, training image frames that are image frames in which instruction related to the task is captured; and generation means for generating training data related to the instruction captured in the image frame sequence based on the training image frames. (Supplementary Note 2) The training data generation device according to Supplementary Note 1, wherein the detection means calculates, for each of a plurality of task types, a probability that the task of that type is captured in the image frames, and detects the image frame as the training image frame if the maximum value of the calculated probabilities is equal to or less than a threshold. (Supplementary Note 3) The training data generation device according to Supplementary Note 1 or 2, wherein the generation means includes the training image frame in the training data. (Supplementary Note 4) The training data generation device according to Supplementary Note 1 or 2, wherein the generation means includes the task type indicated by the training image frame in the training data. (Supplementary Note 5) The teaching data generation device according to Supplementary Note 1 or 2, wherein the detection means detects, from the image frame sequence, a teaching image frame sequence consisting of a plurality of the teaching image frames that are successive in time series, and the generation means includes, in the teaching data, a period of instruction represented by the teaching image frame sequence. (Supplementary Note 6) The teaching data generation device according to Supplementary Note 1 or 2, wherein the acquisition means acquires audio data corresponding to the image frame sequence, the detection means detects, from the image frame sequence, a teaching image frame sequence consisting of a plurality of the teaching image frames that are successive in time series, and the extraction means extracts, from the acquired audio data, audio data corresponding to the teaching image frame sequence as teaching audio data, and the generation means generates the teaching data based on the teaching audio data. (Supplementary Note 7) The teaching data generation device according to Supplementary Note 6, wherein the generation means includes the teaching audio data in the teaching data. (Supplementary Note 8) The teaching data generation device according to Supplementary Note 6, wherein the generation means calculates text representing the content of utterances included in the teaching audio data and includes the calculated text in the teaching data.(Supplementary Note 9) A teaching data generation method executed by a computer, comprising: an acquisition step of acquiring an image frame sequence in which a task is captured, a detection step of detecting, from the image frame sequence, teaching image frames which are image frames in which instruction on the task is captured, and a generation step of generating teaching data on the instruction captured in the image frame sequence based on the teaching image frames. (Supplementary Note 10) A program that causes a computer to execute: an acquisition step of acquiring an image frame sequence in which a task is captured, a detection step of detecting, from the image frame sequence, teaching image frames which are image frames in which instruction on the task is captured, and a generation step of generating teaching data on the instruction captured in the image frame sequence based on the teaching image frames.

[0090] Some or all of the elements (e.g., configurations and functions) described in Supplementary Notes 2 to 8 that are dependent on Supplementary Note 1 may also be dependent on Supplementary Notes 9 and 10 in the same dependency relationship as Supplementary Notes 2 to 8. Some or all of the elements described in any Supplementary Note may be applied to various hardware, software, recording means for recording software, systems, and methods.

[0091] This application claims priority based on Japanese Patent Application No. 2024-110249, filed on July 9, 2024, the disclosure of which is incorporated herein in its entirety by reference.

[0092] REFERENCE SIGNS LIST 10 Image frame sequence 12 Image frame 30 Training image frame 40 Training data 50 Video data 60 Audio data 70 Training audio data 80 Training image frame sequence 1000 Computer 1020 Bus 1040 Processor 1060 Memory 1080 Storage device 1100 Input / output interface 1120 Network interface 2000 Training data generating device 2020 Acquisition unit 2040 Detection unit 2060 Generation unit 2080 Extraction unit

Claims

1. A teaching data generation device having: an acquisition means for acquiring a sequence of image frames in which a task is captured; a detection means for detecting, from the sequence of image frames, teaching image frames which are image frames in which instruction on the task is captured; and a generation means for generating teaching data on the instruction captured in the sequence of image frames based on the teaching image frames.

2. The training data generation device of claim 1, wherein the detection means calculates the probability that a task of a plurality of task types is captured in the image frame, and if the maximum value of the calculated probabilities is equal to or less than a threshold value, detects the image frame as the training image frame.

3. The training data generating device according to claim 1 or 2, wherein said generating means includes said training image frame in said training data.

4. The training data generating device according to claim 1 or 2, wherein said generating means includes the type of work indicated by said training image frame in said training data.

5. A teaching data generating device as described in claim 1 or 2, wherein the detection means detects a teaching image frame sequence consisting of a plurality of teaching image frames that are consecutive in time series from the image frame sequence, and the generation means includes in the teaching data the period of instruction represented by the teaching image frame sequence.

6. A teaching data generation device as described in claim 1 or 2, wherein the acquisition means acquires audio data corresponding to the image frame sequence, the detection means detects a teaching image frame sequence consisting of a plurality of teaching image frames that are consecutive in time series from the image frame sequence, and the device has extraction means for extracting audio data corresponding to the teaching image frame sequence from the acquired audio data as teaching audio data, and the generation means generates the teaching data based on the teaching audio data.

7. The training data generating device according to claim 6, wherein said generating means includes said training voice data in said training data.

8. The training data generating device according to claim 6, wherein said generating means calculates text representing the content of the utterance contained in said training voice data, and includes said calculated text in said training data.

9. A computer-executable teaching data generation method comprising: an acquisition step of acquiring a sequence of image frames in which a task is captured; a detection step of detecting teaching image frames from the sequence of image frames, which are image frames in which instruction regarding the task is captured; and a generation step of generating teaching data regarding the instruction captured in the sequence of image frames based on the teaching image frames.

10. A program that causes a computer to execute the following steps: an acquisition step of acquiring a sequence of image frames in which a task is captured; a detection step of detecting instruction image frames from the sequence of image frames, which are image frames in which instruction regarding the task is captured; and a generation step of generating instruction data regarding the instruction captured in the sequence of image frames based on the instruction image frames.

Citation Information

Patent Citations

  • Video distribution processing system for sports program

    JP2007065958A

  • Server device and instructor supporting system and instructor support method and program

    JP2021068131A