Image processing device and image processing method
The video processing device addresses the inability of existing systems to provide procedure names by detecting action intervals and associating them with corresponding task steps, enhancing user comprehension of video content.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- MITSUBISHI ELECTRIC CORP
- Filing Date
- 2025-06-04
- Publication Date
- 2026-05-07
AI Technical Summary
Existing video processing devices can detect action segments but fail to provide the corresponding procedure names, making it difficult for users to understand the tasks being performed.
A video processing device that includes a procedure group candidate acquisition unit to acquire multiple procedure group candidates, an action interval detection unit to detect action intervals, and a procedure group candidate selection unit to select and present the corresponding procedure names based on feature changes in the video.
Enables users to easily identify the procedure names associated with action intervals, facilitating better understanding and performance of tasks shown in the video.
Smart Images

Figure 0007855153000001 
Figure 0007855153000002 
Figure 0007855153000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to a video processing apparatus and a video processing method.
Background Art
[0002] There is a video processing apparatus that detects an action section, which is a section in which each of a series of procedures for executing a task shown in a video is shown, in the video. As such a video processing apparatus, for example, in Non-Patent Document 1, there is disclosed a video processing apparatus including a feature amount extraction unit that extracts a feature amount of a task shown in a video and outputs time-series data indicating a temporal change of the feature amount, and a clustering unit that detects an action section in which each of a series of procedures is shown based on the temporal change of the feature amount indicated by the time-series data output from the feature amount extraction unit.
Prior Art Documents
Non-Patent Documents
[0003]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] The video processing device disclosed in Non-Patent Document 1 can detect action segments in which each step of a series of procedures is shown, but it has the problem of not being able to obtain the procedure name of the procedure corresponding to the action segment. As a result, even if a user looks at the detection results of the video processing device, they may not be able to easily understand what procedure corresponds to the action segment.
[0005] This disclosure was made to solve the above-mentioned problems and aims to provide a video processing device that can obtain the procedure name of a procedure corresponding to an action interval. [Means for solving the problem]
[0006] The video processing device according to this disclosure includes a procedure group candidate acquisition unit that acquires multiple procedure group candidates, each including multiple possible steps for executing a task shown in a video and the name of each step in those steps; and an action interval detection unit that detects action intervals in a video, which are sections in a video in which each of the steps for executing a task may be shown, based on the change in the feature quantities of the task shown in the video over time. The video processing device also includes a procedure group candidate selection unit that selects a procedure group candidate from among the multiple procedure group candidates acquired by the procedure group candidate acquisition unit that corresponds to an action interval detected by the action interval detection unit. [Effects of the Invention]
[0007] According to this disclosure, the procedure name of the procedure corresponding to the action interval can be obtained. [Brief explanation of the drawing]
[0008] [Figure 1] This is a configuration diagram showing the video processing device according to Embodiment 1. [Figure 2] This is a hardware configuration diagram showing the hardware of the video processing device according to Embodiment 1. [Figure 3]This is a hardware configuration diagram of a computer when the video processing device is implemented using software or firmware. [Figure 4] This is a flowchart showing the image processing method, which is the processing procedure of an image processing device. [Figure 5] This is an explanatory diagram showing an example of three candidate procedure sets created by generative AI. [Figure 6] This is an explanatory diagram illustrating an example where multiple interval group candidates are output from a neural model such as a TAS model. [Figure 7] This is an action send date graph showing an example of interval group candidate selection by the procedure group candidate selection processing unit 3b. [Figure 8] This is an explanatory diagram showing an example of an action interval and procedure name presented by the presentation processing unit 4. [Figure 9] This is a configuration diagram showing an image processing device according to Embodiment 2. [Figure 10] This is a hardware configuration diagram showing the hardware of the video processing device according to Embodiment 2. [Figure 11] This is a configuration diagram showing the video processing device according to Embodiment 3. [Figure 12] This is a hardware configuration diagram showing the hardware of the video processing device according to Embodiment 3. [Figure 13] This is a flowchart showing the image processing method, which is the processing procedure of an image processing device. [Modes for carrying out the invention]
[0009] To provide a more detailed explanation of this disclosure, the forms for implementing this disclosure will be described below with reference to the attached drawings.
[0010] Embodiment 1. Figure 1 is a configuration diagram showing an image processing device according to Embodiment 1. Figure 2 is a hardware configuration diagram showing the hardware of the video processing device according to Embodiment 1. The video processing apparatus shown in FIG. 1 includes a procedure group candidate acquisition unit 1, an action section detection unit 2, a procedure group candidate selection unit 3, and a presentation processing unit 4.
[0011] The procedure group candidate acquisition unit 1 is realized, for example, by a procedure group candidate acquisition circuit 21 shown in FIG. 2. The procedure group candidate acquisition unit 1 acquires a plurality of different procedure group candidates. A procedure group candidate includes a plurality of possible procedures for a series of procedures for executing a task shown in a video and the respective procedure names in the plurality of procedures. The procedure group candidate acquisition unit 1 outputs information indicating a plurality of procedure group candidates to the procedure group candidate selection unit 3.
[0012] The action section detection unit 2 is realized, for example, by an action section detection circuit 22 shown in FIG. 2. The action section detection unit 2 includes a feature amount extraction unit 2a and a section group candidate acquisition unit 2b. The action section detection unit 2 detects an action section, which is a section in the video where each of a series of procedures for executing a task may be shown, based on the change over time of the feature amount of the task shown in the video. The action section detection unit 2 outputs the detection result of the action section to the procedure group candidate selection unit 3.
[0013] The feature amount extraction unit 2a acquires a video in which a task is shown. The feature amount extraction unit 2a extracts the feature amount of the task shown in the video. The feature amount extraction unit 2a outputs time-series data indicating the time change of the feature amount to the section group candidate acquisition unit 2b.
[0014] The section group candidate acquisition unit 2b acquires time-series data from the feature amount extraction unit 2a. The section group candidate acquisition unit 2b acquires a plurality of different section group candidates based on the time change of the feature amount indicated by the time-series data. Candidate interval groups include action intervals that may represent each of the steps in a sequence of steps required to perform a task. The interval group candidate acquisition unit 2b outputs information indicating multiple interval group candidates to the procedure group candidate selection unit 3 as a result of detecting action intervals.
[0015] The procedure group candidate selection unit 3 is implemented, for example, by the procedure group candidate selection circuit 23 shown in Figure 2. The procedure group candidate selection unit 3 includes a matching processing unit 3a and a procedure group candidate selection processing unit 3b. The procedure group candidate selection unit 3 obtains information indicating multiple procedure group candidates from the procedure group candidate acquisition unit 1 and obtains the action interval detection result from the action interval detection unit 2. The procedure group candidate selection unit 3 selects a procedure group candidate from among the multiple procedure group candidates acquired by the procedure group candidate acquisition unit 1 that corresponds to the action interval detected by the action interval detection unit 2. The procedure group candidate selection unit 3 outputs information indicating the selected procedure group candidate to the presentation processing unit 4.
[0016] The matching processing unit 3a obtains information indicating multiple procedure group candidates from the procedure group candidate acquisition unit 1, and obtains information indicating multiple interval group candidates from the interval group candidate acquisition unit 2b. The matching processing unit 3a compares each candidate procedure group with each candidate interval group. The matching processing unit 3a outputs the matching result between each candidate procedure group and each candidate interval group to the candidate procedure group selection processing unit 3b.
[0017] The procedure group candidate selection processing unit 3b obtains information indicating multiple procedure group candidates from the procedure group candidate acquisition unit 1, and obtains information indicating multiple interval group candidates from the interval group candidate acquisition unit 2b. Furthermore, the procedure group candidate selection processing unit 3b obtains the matching result from the matching processing unit 3a. The procedure group candidate selection processing unit 3b selects, based on the matching result of the matching processing unit 3a, a group group candidate that includes an action group that shows each of the steps in a series of steps for executing a task, from among the multiple group group candidates acquired by the group group candidate acquisition unit 2b. The procedure group candidate selection processing unit 3b selects a procedure group candidate that corresponds to an action interval included in the selected interval group candidate from among the multiple procedure group candidates acquired by the procedure group candidate acquisition unit 1, based on the matching result of the matching processing unit 3a. The procedure group candidate selection processing unit 3b outputs information indicating the selected interval group candidate and information indicating the selected procedure group candidate to the presentation processing unit 4.
[0018] The presentation processing unit 4 is implemented, for example, by the presentation processing circuit 24 shown in Figure 2. The presentation processing unit 4 obtains information indicating the interval group candidate and information indicating the procedure group candidate from the procedure group candidate selection processing unit 3b. The presentation processing unit 4 presents the action intervals included in the candidate interval group selected by the procedure group candidate processing unit 3b, and the procedure names of the procedures included in the candidate interval group selected by the procedure group candidate processing unit 3b. Specifically, the presentation processing unit 4 displays the action interval and procedure name on a display device (not shown), for example.
[0019] In Figure 1, the components of the video processing device—the procedure group candidate acquisition unit 1, the action interval detection unit 2, the procedure group candidate selection unit 3, and the presentation processing unit 4—are assumed to be implemented by dedicated hardware as shown in Figure 2. Specifically, the video processing device is assumed to be implemented by a procedure group candidate acquisition circuit 21, an action interval detection circuit 22, a procedure group candidate selection circuit 23, and a presentation processing circuit 24. Each of the procedure group candidate acquisition circuit 21, action interval detection circuit 22, procedure group candidate selection circuit 23, and presentation processing circuit 24 can be, for example, a single circuit, a composite circuit, a programmed processor, a parallel programmed processor, an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), or a combination thereof.
[0020] The components of the video processing device are not limited to those implemented by dedicated hardware; the video processing device may also be implemented by software, firmware, or a combination of software and firmware. Software or firmware is stored as a program in the computer's memory. A computer refers to the hardware that executes programs, and includes, for example, a CPU (Central Processing Unit), GPU (Graphics Processing Unit), central processing unit, processing unit, arithmetic unit, microprocessor, microcomputer, processor, or DSP (Digital Signal Processor).
[0021] Figure 3 is a hardware configuration diagram of a computer when the video processing device is implemented by software or firmware. When the video processing device is implemented using software or firmware, a program is stored in memory 31 that causes the computer to execute each of the processing procedures in the procedure group candidate acquisition unit 1, the action interval detection unit 2, the procedure group candidate selection unit 3, and the presentation processing unit 4. The computer's processor 32 then executes the program stored in memory 31.
[0022] Furthermore, Figure 2 shows an example where each component of the video processing device is implemented by dedicated hardware, and Figure 3 shows an example where the video processing device is implemented by software or firmware, etc. However, this is only one example, and some components of the video processing device may be implemented by dedicated hardware, while the remaining components may be implemented by software or firmware, etc.
[0023] Next, we will explain the operation of the video processing device shown in Figure 1. Figure 4 is a flowchart showing the image processing method, which is the processing procedure of the image processing device. The procedure group candidate acquisition unit 1 acquires multiple procedure group candidates that are different from each other (step ST1 in Figure 4). Specifically, the procedure group candidate acquisition unit 1 acquires a request from an external source to create a series of procedures for executing the tasks shown in the video. Then, the procedure group candidate acquisition unit 1 provides a creation request to the generation AI (Artificial Intelligence) and acquires multiple procedure group candidates from the generation AI. The procedure group candidate acquisition unit 1 outputs information indicating multiple procedure group candidates to the procedure group candidate selection unit 3. The creator of a request to create a series of steps could be, for example, the creator of the video, a third party other than the video creator, or a user who wants to check the content of the tasks shown in the video. Furthermore, if the creation requests corresponding to each of the multiple tasks are stored, for example, in a table (not shown), the procedure group candidate acquisition unit 1 may, instead of acquiring the creation requests from an external source, acquire the identification information of the tasks shown in the video and read the creation requests corresponding to the tasks indicated by the identification information from the table.
[0024] A task, for example, means work or an assignment, and the video shows the work or assignment being carried out. If the objective of the task shown in the video is, for example, "to make coffee," then a request to create a series of steps might be, "Please provide three possible steps for making coffee."
[0025] When the generating AI receives a request from the procedure group candidate acquisition unit 1 to create a series of procedures, it creates multiple procedure group candidates corresponding to the creation request and outputs the multiple procedure group candidates to the procedure group candidate acquisition unit 1. When the generating AI is given a request to create a series of steps, for example, "Please give me three possible steps for making coffee," it will create three possible sets of steps (1) to (3) as shown below.
[0026] Candidate procedure group (1) Pick up the cup → Pour in the coffee → Pour in the milk → Stir the coffee Candidate procedure group (2) Pick up a cup → Pour in coffee → Pour in water → Stir the coffee Candidate procedure group (3) Pick up a cup → Pour in coffee → Pour in milk → Pour in sugar → Stir in the coffee
[0027] Figure 5 is an explanatory diagram showing an example of three candidate procedure sets created by generative AI. In Figure 5, the circles indicate procedures, and the topmost candidate set of procedures in the figure corresponds to candidate set of procedures (1) which includes four procedures. In the diagram, the second candidate set of procedures from the top corresponds to candidate set of procedures (2) which contains four procedures. In the diagram, the third candidate set of procedures from the top corresponds to candidate set of procedures (3) which includes five procedures.
[0028] Each of the three candidate procedure groups (1) to (3) includes the name of each procedure it contains. For example, in the case of candidate procedure group (1), a possible procedure name for the procedure "take the cup" could be "cup removal step." A possible procedure name for the procedure "pour the coffee" could be "coffee pouring step." For example, the step name for the procedure "pouring milk" could be "milk pouring step." For example, the step name for the procedure "stirring coffee" could be "stirring step."
[0029] Here, we show an example where each of the candidate procedure groups (1) to (3) contains information indicating the procedure name in addition to multiple procedures. However, this is just one example, and the procedures contained within each of the candidate procedure groups (1) to (3) may also serve as information indicating the procedure name. In the video processing device shown in Figure 1, the procedure group candidate acquisition unit 1 provides the generation AI with a request to create a series of procedures for executing a task, and acquires multiple procedure group candidates from the generation AI. However, this is just one example, and the procedure group candidate acquisition unit 1 may also receive multiple procedure group candidates from outside the video processing device, for example, via a network.
[0030] The action interval detection unit 2 acquires a video showing the task. The action interval detection unit 2 detects action intervals in the video, which are segments in the video that may contain each of the steps required to perform the task, based on how the feature quantities of the task shown in the video change over time (step ST2 in Figure 4). The action interval detection unit 2 outputs the detection result of the action interval to the procedure group candidate selection unit 3. As a method for detecting action intervals by the action interval detection unit 2, for example, a well-known technique called TAS (Temporal Action Segmentation) can be used. The following describes in detail the action interval detection process performed by the action interval detection unit 2.
[0031] The feature extraction unit 2a acquires a video in which the task is shown. The feature extraction unit 2a extracts the features of the task shown in the video. Since the feature extraction process itself is a well-known technique, a detailed explanation will be omitted. The feature extraction unit 2a outputs time-series data showing the change in features over time to the interval group candidate acquisition unit 2b.
[0032] The interval group candidate acquisition unit 2b acquires time series data from the feature extraction unit 2a. The interval group candidate acquisition unit 2b acquires multiple distinct interval group candidates based on the time changes of features shown in the time series data, for example, using a technique called TAS. The interval group candidate acquisition unit 2b outputs information indicating multiple interval group candidates to the procedure group candidate selection unit 3 as a result of detecting action intervals.
[0033] In the video processing device shown in Figure 1, the feature extraction unit 2a extracts the features of the task, and the interval group candidate acquisition unit 2b acquires multiple interval group candidates based on the time changes of the features shown in the time series data. However, this is only one example, and the interval group candidate acquisition unit 2b may also be configured to acquire multiple interval group candidates from a pre-trained model that has learned the correspondence between the features of the task and the action intervals, by providing the time series data output from the feature extraction unit 2a to the pre-trained model. Examples of pre-trained models include existing neural models such as TAS models.
[0034] Figure 6 is an explanatory diagram illustrating an example where multiple interval group candidates are output from a neural model such as a TAS model. In Figure 6, four different interval group candidates, IGC1 to IGC4, are shown as examples. Specifically, examples include IGC1, a candidate interval group including action intervals (1), (3), and (5)-(7); IGC2, a candidate interval group including action intervals (1), (4), and (5)-(7); IGC3, a candidate interval group including action intervals (2), (3), and (5)-(7); and IGC4, a candidate interval group including action intervals (2), (4), and (5)-(7).
[0035] The procedure group candidate selection unit 3 obtains information indicating multiple procedure group candidates from the procedure group candidate acquisition unit 1 and obtains the action interval detection result from the action interval detection unit 2. The procedure group candidate selection unit 3 selects a procedure group candidate from among the multiple procedure group candidates acquired by the procedure group candidate acquisition unit 1 that corresponds to the action interval detected by the action interval detection unit 2 (step ST3 in Figure 4). The procedure group candidate selection unit 3 outputs information indicating the selected procedure group candidate to the presentation processing unit 4. The following describes in detail the process by which the procedure group candidate selection unit 3 selects procedure group candidates.
[0036] The matching processing unit 3a obtains information from the procedure group candidate acquisition unit 1 indicating multiple procedure group candidates, for example, information indicating procedure group candidates (1) to (3). The matching processing unit 3a obtains information from the interval group candidate acquisition unit 2b indicating multiple interval group candidates, for example, information indicating four different interval group candidates IGC1 to IGC4. The matching processing unit 3a converts information indicating candidate procedure groups (1) to (3) into information in a certain feature space, and converts information indicating candidate interval groups IGC1 to IGC4 into information in the same feature space as the above feature space. The process of converting this information into feature space information is a well-known technique, so a detailed explanation will be omitted.
[0037] The matching processing unit 3a compares each of the procedure group candidates (1) to (3) with each of the interval group candidates IGC1 to IGC4 in the feature space described above. Specifically, the matching processing unit 3a calculates the similarity between each of the candidate procedure groups (1) to (3) and each of the candidate interval groups IGC1 to IGC4 in the feature space described above. The process of calculating the similarity between candidate procedure groups and candidate interval groups is a well-known technique, so a detailed explanation will be omitted. The matching processing unit 3a outputs the similarity calculation result as the matching result to the procedure group candidate selection processing unit 3b.
[0038] The procedure group candidate selection processing unit 3b acquires information from the procedure group candidate acquisition unit 1 that indicates multiple procedure group candidates, for example, information indicating procedure group candidates (1) to (3). The procedure group candidate selection processing unit 3b obtains information from the interval group candidate acquisition unit 2b indicating four possible interval group candidates, IGC1 to IGC4, as information indicating multiple interval group candidates. The procedure group candidate selection processing unit 3b obtains the similarity calculation result as the matching result from the matching processing unit 3a. The procedure group candidate selection processing unit 3b selects the combination of procedure group candidate (1) to (3) and interval group candidate IGC1 to IGC4 that has the highest similarity, based on the matching results of the matching processing unit 3a. The procedure group candidate selection processing unit 3b selects interval group candidate IGC1 from among the four interval group candidates IGC1 to IGC4 if the combination of procedure group candidate (3) and interval group candidate IGC1 has the highest similarity.
[0039] Figure 7 is an action senddate graph showing an example of interval group candidate selection by the procedure group candidate selection processing unit 3b. Figure 7 shows an example where the candidate interval group IGC1 is selected. In Figure 7, the horizontal axis represents time, and the vertical axis represents the posterior probability. The procedure group candidate selection processing unit 3b outputs information indicating the selected interval group candidate AGC1 and information indicating the selected procedure group candidate (3) to the presentation processing unit 4.
[0040] In the image processing device shown in Figure 1, the procedure group candidate selection processing unit 3b selects the combination of procedure group candidate and interval group candidate that has the highest similarity. The combination of procedure group candidate and interval group candidate is not limited to the combination with the highest similarity; the procedure group candidate selection processing unit 3b may, within a range that does not pose a practical problem, select, for example, the second-highest similar combination or the third-highest similar combination.
[0041] The presentation processing unit 4 obtains information indicating the selected interval group candidate and information indicating the selected procedure group candidate from the procedure group candidate selection processing unit 3b. The presentation processing unit 4 obtains, for example, information indicating the interval group candidate IGC1 and information indicating the procedure group candidate (3). The presentation processing unit 4 displays the action intervals included in the candidate interval group indicated by the acquired information on a display device (not shown) (step ST4 in Figure 4). Specifically, if the interval group candidate indicated by the acquired information is interval group candidate IGC1, the presentation processing unit 4 displays action intervals (1), action interval (3), and action intervals (5) to (7) as action intervals included in interval group candidate IGC1, for example, on a display device (not shown), as shown in Figure 8. Furthermore, the presentation processing unit 4 displays the names of the procedures included in the candidate procedure group indicated by the acquired information on a display device (not shown) (step ST4 in Figure 4). Specifically, if the procedure group candidate indicated by the acquired information is procedure group candidate (3), the presentation processing unit 4 displays the procedure names "cup removal step," "coffee pouring step," "milk pouring step," "sugar pouring step," and "stirring step" on a display device (not shown), for example, as shown in Figure 8.
[0042] Figure 8 is an explanatory diagram showing an example of an action interval and procedure name presented by the presentation processing unit 4. Users who view the display on the device can easily recognize the tasks shown in the video because they can see the names of the steps involved. Therefore, users who view the display on the device can easily perform the tasks shown in the video.
[0043] In the above embodiment 1, the video processing device is configured to include a procedure group candidate acquisition unit 1 that acquires multiple procedure group candidates, each containing multiple possible steps for executing a task shown in a video and the name of each step in those steps; and an action section detection unit 2 that detects action sections in a video, which are sections in the video in which each of the steps for executing a task may be shown, based on the change in the feature quantities of the task shown in the video over time. The video processing device also includes a procedure group candidate selection unit 3 that selects a procedure group candidate from among the multiple procedure group candidates acquired by the procedure group candidate acquisition unit 1 that corresponds to an action section detected by the action section detection unit 2. Therefore, the video processing device can acquire the name of the step corresponding to the action section.
[0044] In Embodiment 1, the video processing device is configured such that the action interval detection unit 2 includes a feature extraction unit 2a that extracts feature quantities of a task shown in the video and outputs time-series data showing the temporal changes in these feature quantities, and an interval group candidate acquisition unit 2b that acquires multiple interval group candidates, each containing an action interval that may represent a set of steps for executing a task, based on the temporal changes in feature quantities shown by the time-series data output from the feature extraction unit 2a. Furthermore, the video processing device is configured such that the procedure group candidate selection unit 3 includes a matching processing unit 3a that compares each procedure group candidate acquired by the procedure group candidate acquisition unit 1 with each interval group candidate acquired by the interval group candidate acquisition unit 2b, and a procedure group candidate selection processing unit 3b that, based on the matching results of the matching processing unit 3a, selects an interval group candidate from among the multiple interval group candidates acquired by the interval group candidate acquisition unit 2b that contains an action interval that represents a set of steps for executing a task, and also selects a procedure group candidate from among the multiple procedure group candidates acquired by the procedure group candidate acquisition unit 1 that corresponds to an action interval included in the selected interval group candidate. Therefore, the video processing device can select a group of candidate intervals containing action intervals in which each of the steps for performing a task may be shown, and a group of candidate steps corresponding to action intervals in which each of the steps may be shown.
[0045] In Embodiment 1, the video processing device is configured such that the action interval detection unit 2 includes a feature extraction unit 2a that extracts feature quantities of tasks shown in the video and outputs time-series data showing the temporal changes in these feature quantities, and an interval group candidate acquisition unit 2b that provides the time-series data output from the feature extraction unit 2a to a trained model that has learned the correspondence between task feature quantities and action intervals, and acquires from the trained model a plurality of interval group candidates that include action intervals in which each of the steps for executing a task may be shown. Furthermore, the video processing device is configured such that the procedure group candidate selection unit 3 includes a matching processing unit 3a that matches each procedure group candidate acquired by the procedure group candidate acquisition unit 1 with each section group candidate acquired by the section group candidate acquisition unit 2b, and a procedure group candidate selection processing unit 3b that, based on the matching result of the matching processing unit 3a, selects a section group candidate from among a plurality of section group candidates acquired by the section group candidate acquisition unit 2b that contains an action section showing each of the steps in a series of steps for executing a task, and also selects a procedure group candidate from among a plurality of procedure group candidates acquired by the procedure group candidate acquisition unit 1 that corresponds to an action section included in the selected section group candidate. Therefore, the video processing device can select a section group candidate that contains an action section that may show each of the steps in a series of steps for executing a task, and a procedure group candidate that corresponds to an action section that may show each of the steps in a series of steps.
[0046] In Embodiment 1, the video processing device is configured such that the matching processing unit 3a calculates the similarity between each procedure group candidate acquired by the procedure group candidate acquisition unit 1 and each interval group candidate acquired by the interval group candidate acquisition unit 2b, and outputs the similarity calculation result as the matching result to the procedure group candidate selection processing unit 3b. Therefore, the video processing device can select a combination of procedure group candidates and interval group candidates with a high similarity.
[0047] In Embodiment 1, the video processing device is configured to include a presentation processing unit 4 that presents the action intervals included in the candidate interval group selected by the candidate interval group processing unit 3b, and the procedure names of the procedures included in the candidate interval group selected by the candidate interval group processing unit 3b. Therefore, the video processing device can present the action intervals and procedure names.
[0048] In Embodiment 1, the video processing device is configured such that the procedure group candidate acquisition unit 1 provides the generation AI with a request to create a series of procedures for executing a task shown in the video, and acquires multiple procedure group candidates from the generation AI. Therefore, the video processing device can acquire multiple procedure group candidates simply by providing the generation AI with a request to create a series of procedures.
[0049] In the image processing device shown in Figure 1, the matching processing unit 3a converts information representing candidate procedure groups (1) to (3) into information in a certain feature space, and converts information representing candidate interval groups IGC1 to IGC4 into information in the same feature space as the above feature space. However, this is just one example, and the matching processing unit 3a may, for example, convert the information representing candidate interval groups IGC1 to IGC4 into the feature space of the information representing candidate procedure groups (1) to (3) without converting the information representing candidate procedure groups (1) to (3).
[0050] Embodiment 2. Embodiment 2 describes a video processing device in which a procedure group candidate acquisition unit 5 acquires multiple procedure group candidates based on text indicating the purpose of the task shown in the video.
[0051] Figure 9 is a configuration diagram showing an image processing device according to Embodiment 2. In Figure 9, the same reference numerals as in Figure 1 indicate the same or corresponding parts, so a detailed explanation is omitted. Figure 10 is a hardware configuration diagram showing the hardware of the video processing device according to Embodiment 2. In Figure 10, the same reference numerals as in Figure 2 indicate the same or corresponding parts, so a detailed explanation is omitted. The video processing device shown in Figure 9 includes a procedure group candidate acquisition unit 5, an action interval detection unit 2, a procedure group candidate selection unit 3, and a presentation processing unit 4.
[0052] The procedure group candidate acquisition unit 5 is implemented, for example, by the procedure group candidate acquisition circuit 25 shown in Figure 10. The procedure group candidate acquisition unit 5 acquires multiple procedure group candidates based on the text indicating the purpose of the task shown in the video and the action labels. The action labels are labels that indicate the action interval, which is the section in which each of the series of steps for executing the task is shown, and the name of the step corresponding to that action interval. The procedure group candidate acquisition unit 5 outputs information indicating multiple procedure group candidates to the procedure group candidate selection unit 3.
[0053] In Figure 9, it is assumed that the components of the video processing device—the procedure group candidate acquisition unit 5, the action interval detection unit 2, the procedure group candidate selection unit 3, and the presentation processing unit 4—are each implemented by dedicated hardware as shown in Figure 10. Specifically, it is assumed that the video processing device is implemented by a procedure group candidate acquisition circuit 25, an action interval detection circuit 22, a procedure group candidate selection circuit 23, and a presentation processing circuit 24. Each of the procedure group candidate acquisition circuit 25, action interval detection circuit 22, procedure group candidate selection circuit 23, and presentation processing circuit 24 can be, for example, a single circuit, a composite circuit, a programmed processor, a parallel programmed processor, an ASIC, an FPGA, or a combination thereof.
[0054] The components of the video processing device are not limited to those implemented by dedicated hardware; the video processing device may also be implemented by software, firmware, or a combination of software and firmware. When the video processing device is implemented using software or firmware, a program is stored in the memory 31 shown in Figure 3 that causes the computer to execute the respective processing procedures in the procedure group candidate acquisition unit 5, the action interval detection unit 2, the procedure group candidate selection unit 3, and the presentation processing unit 4. Then, the processor 32 shown in Figure 3 executes the program stored in the memory 31.
[0055] Furthermore, Figure 10 shows an example where each component of the video processing device is implemented by dedicated hardware, and Figure 3 shows an example where the video processing device is implemented by software or firmware, etc. However, this is only one example, and some components of the video processing device may be implemented by dedicated hardware, while the remaining components may be implemented by software or firmware, etc.
[0056] Next, the operation of the video processing device shown in Figure 9 will be explained. However, since all parts except the procedure group candidate acquisition unit 5 are the same as those in the video processing device shown in Figure 1, only the operation of the procedure group candidate acquisition unit 5 will be explained here.
[0057] The procedure group candidate acquisition unit 5 acquires text indicating the purpose of the task shown in the video, as well as action labels. If the objective of the task shown in the video is, for example, "to make coffee," then the text indicating the objective of the task would be "to make coffee." Action labels indicate the action interval, which shows each step in a series of steps for performing a task, and the name of the step corresponding to that action interval. The procedure group candidate acquisition unit 5 acquires multiple procedure group candidates based on text indicating the task's objective and action labels.
[0058] Specifically, the procedure group candidate acquisition unit 5 provides a language model such as an LLM (Large Language Model) with text indicating the purpose of the task shown in the video and an action label, and acquires multiple procedure group candidates from the language model. The procedure group candidates include multiple possible steps for a series of steps to perform the task shown in the video, and the names of each step in those steps. During training, the language model learns candidate sets of procedures that correspond to the task's objective. In other words, during training, the language model learns the correspondence between the task's objective, the action intervals that represent each of the steps necessary to perform the task, and the procedure names corresponding to those action intervals. During inference, when the language model receives text and labels indicating the purpose of the task shown in the video from the procedure candidate acquisition unit 5, it outputs multiple procedure candidate sets to the procedure candidate acquisition unit 5. The procedure group candidate acquisition unit 5 outputs information indicating multiple procedure group candidates to the procedure group candidate selection unit 3.
[0059] The video processing device shown in Figure 9 demonstrates a language model that, during training, learns the correspondence between the task's objective, the action segments in which each of the steps for executing the task is shown, and the step names corresponding to those action segments. However, this is merely one example, and the language model may not have learned the correspondence between the task's objective, the action segments, and the step names corresponding to those action segments. In this case, during inference, when the language model receives text indicating the task's objective and action labels shown in the video from the procedure group candidate acquisition unit 5, it outputs multiple procedure group candidates to the procedure group candidate acquisition unit 5 based on its common sense knowledge.
[0060] In the above embodiment 2, the video processing device is configured such that the procedure group candidate acquisition unit 5 acquires multiple procedure group candidates based on text indicating the purpose of the task shown in the video, and labels indicating action sections, which are sections in which each of the series of steps for executing the task is shown, and the procedure names of the steps corresponding to those action sections. Therefore, the video processing device can acquire multiple procedure group candidates simply by providing the text and labels to, for example, a language model.
[0061] Embodiment 3. Embodiment 3 describes an image processing apparatus comprising a feature extraction unit 6, a procedure name acquisition unit 7, a procedure name modification unit 8, and a presentation processing unit 4.
[0062] Figure 11 is a configuration diagram showing an image processing device according to Embodiment 3. In Figure 11, the same reference numerals as in Figure 1 indicate the same or corresponding parts, so a detailed explanation is omitted. Figure 12 is a hardware configuration diagram showing the hardware of the video processing device according to Embodiment 3. In Figure 12, the same reference numerals as in Figure 2 indicate the same or corresponding parts, so a detailed explanation is omitted. The image processing device shown in Figure 11 includes a feature extraction unit 6, a procedure name acquisition unit 7, a procedure name modification unit 8, and a presentation processing unit 4.
[0063] The feature extraction unit 6 is implemented, for example, by the feature extraction circuit 26 shown in Figure 12. The feature extraction unit 6 acquires a video in which the task is shown. The feature extraction unit 6 extracts the features of the task shown in the video. The feature extraction unit 6 outputs time-series data showing the change in features over time to the procedure name acquisition unit 7.
[0064] The procedure name acquisition unit 7 is implemented, for example, by the procedure name acquisition circuit 27 shown in Figure 12. The procedure name acquisition unit 7 acquires time-series data showing the change in features over time from the feature extraction unit 6. The procedure name acquisition unit 7 provides the time-series data output from the feature extraction unit 6 to a trained model 9 that has learned the correspondence between the features of a task and the procedure names of the steps for executing the task, and acquires the procedure names of the series of steps for executing the task shown in the video from the trained model 9. The procedure name acquisition unit 7 outputs information indicating the procedure name to the procedure name modification unit 8.
[0065] During training, the trained model 9, given time-series data showing the temporal changes in the features of the tasks shown in the video, along with action labels, learns the correspondence between the features of the tasks shown in the video, the action intervals, and the procedure names. During inference, when the trained model 9 is given time-series data output from the feature extraction unit 6 via the procedure name acquisition unit 7, it outputs to the procedure name acquisition unit 7 the action interval, which is the interval in which each of the steps for executing the task from which features are extracted by the feature extraction unit 6 is shown, and the procedure name of the step corresponding to the action interval. In the image processing device shown in Figure 11, the trained model 9 is located outside the image processing device. However, this is just one example, and the trained model 9 may also be located inside the image processing device.
[0066] The procedure name correction unit 8 is implemented, for example, by the procedure name correction circuit 28 shown in Figure 12. The procedure name modification unit 8 obtains the procedure names of each procedure from the procedure name acquisition unit 7. The procedure name modification unit 8 modifies the procedure name obtained by the procedure name acquisition unit 7 using action labels that indicate the procedure name of the procedure corresponding to the action interval, which is the interval in which each of the series of procedures for executing the task from which features are extracted by the feature extraction unit 6 is shown. The procedure name modification unit 8 outputs information indicating the modified procedure name to the presentation processing unit 4.
[0067] In Figure 11, the feature extraction unit 6, procedure name acquisition unit 7, procedure name modification unit 8, and presentation processing unit 4, which are components of the video processing device, are assumed to be implemented by dedicated hardware as shown in Figure 12. That is, the video processing device is assumed to be implemented by a feature extraction circuit 26, a procedure name acquisition circuit 27, a procedure name modification circuit 28, and a presentation processing circuit 24. Each of the feature extraction circuit 26, procedure name acquisition circuit 27, procedure name correction circuit 28, and presentation processing circuit 24 can be, for example, a single circuit, a composite circuit, a programmed processor, a parallel programmed processor, an ASIC, an FPGA, or a combination thereof.
[0068] The components of the video processing device are not limited to those implemented by dedicated hardware; the video processing device may also be implemented by software, firmware, or a combination of software and firmware. When the video processing device is implemented using software or firmware, a program is stored in the memory 31 shown in Figure 3 that causes the computer to execute the respective processing procedures in the feature extraction unit 6, the procedure name acquisition unit 7, the procedure name modification unit 8, and the presentation processing unit 4. Then, the processor 32 shown in Figure 3 executes the program stored in the memory 31.
[0069] Furthermore, Figure 12 shows an example in which each component of the video processing device is implemented by dedicated hardware, and Figure 3 shows an example in which the video processing device is implemented by software or firmware, etc. However, this is only one example, and some components of the video processing device may be implemented by dedicated hardware, while the remaining components may be implemented by software or firmware, etc.
[0070] Next, we will explain the operation of the video processing device shown in Figure 11. Figure 13 is a flowchart showing the image processing method, which is the processing procedure of the image processing device. The feature extraction unit 6 acquires a video in which the task is shown. The feature extraction unit 6 extracts the features of the task shown in the video, similar to the feature extraction unit 2a shown in Figure 1 (step ST11 in Figure 13). The feature extraction unit 6 outputs time-series data showing the change in features over time to the procedure name acquisition unit 7.
[0071] The procedure name acquisition unit 7 acquires time-series data showing the change in features over time from the feature extraction unit 6. The procedure name acquisition unit 7 provides time-series data to the trained model 9 and obtains from the trained model 9 action intervals, which are segments representing each of the steps in a series of steps for executing a task in which features are extracted by the feature extraction unit 6, and the procedure names of the steps corresponding to those action intervals (step ST12 in Figure 13). The procedure name acquisition unit 7 outputs action interval information indicating the action interval and procedure name information indicating the procedure name to the procedure name modification unit 8.
[0072] The procedure name modification unit 8 obtains action interval information and procedure name information from the procedure name acquisition unit 7. The procedure name modification unit 8 obtains an action label that indicates the procedure name of the step corresponding to the action section, in addition to the action section that shows each of the steps in the sequence for executing the task. The procedure name modification unit 8 modifies the procedure name obtained by the procedure name acquisition unit 7 using the action label (step ST13 in Figure 13). The following describes in detail the process of correcting the procedure name by the procedure name correction unit 8.
[0073] Specifically, the procedure name modification unit 8 compares multiple action intervals indicated by the action interval information (hereinafter referred to as "acquired action intervals") with multiple action intervals indicated by the action labels (hereinafter referred to as "labeled action intervals") and identifies identical action intervals. The procedure name modification unit 8 compares the procedure name of the procedure corresponding to the acquired action interval (hereinafter referred to as the "acquired procedure name") with the procedure name indicated by the action label corresponding to the label action interval, which is the same action interval as the acquired action interval (hereinafter referred to as the "labeled procedure name"). If the acquired procedure name and the labeled procedure name match, the procedure name modification unit 8 outputs action interval information indicating the acquired action interval and procedure name information indicating the acquired procedure name to the presentation processing unit 4 without modifying the acquired procedure name. If the acquired procedure name and the label procedure name do not match, the procedure name correction unit 8 corrects the acquired procedure name to the label procedure name and outputs action interval information indicating the acquired action interval and procedure name information indicating the corrected procedure name to the presentation processing unit 4.
[0074] The presentation processing unit 4 obtains action interval information and procedure name information from the procedure name modification unit 8. The display processing unit 4 displays the action interval indicated by the action interval information and the procedure name indicated by the procedure name information on a display device (not shown) (step ST14 in Figure 13).
[0075] In the above embodiment 3, the video processing device is configured to include a feature extraction unit 6 that extracts features of tasks shown in a video and outputs time-series data showing the temporal changes in these features, and a procedure name acquisition unit 7 that provides the time-series data output from the feature extraction unit 6 to a trained model 9 and obtains the procedure names of a series of steps for executing the tasks shown in the video from the trained model 9. Therefore, the video processing device can obtain the procedure names of steps corresponding to action intervals.
[0076] In Embodiment 3, the video processing device is configured to include a procedure name correction unit 8 that corrects the procedure names obtained by the procedure name acquisition unit 7 using labels indicating the procedure names of the procedures corresponding to the action intervals, which are sections in which each of the series of steps for executing a task from which features are extracted by the feature extraction unit 6 is shown. Therefore, the video processing device can optimize the procedure names.
[0077] In the video processing device shown in Figure 11, the procedure name acquisition unit 7 provides time-series data to a trained model 9 and obtains action intervals and procedure names from the trained model 9. However, this is just one example; the procedure name acquisition unit 7 also obtains labels indicating action intervals and the procedure names of the steps corresponding to those action intervals from supplementary materials such as work manuals, standard procedure documents, or recipes. Specifically, the procedure name acquisition unit 7 provides the supplementary materials to a first language model and obtains labels from the first language model. The procedure name acquisition unit 7 may also provide the second language model with text and labels indicating the purpose of the task shown in the video, and obtain multiple procedure group candidates from the second language model.
[0078] The video processing device shown in Figure 11 indicates that if the acquired procedure name and the labeled procedure name do not match, the procedure name correction unit 8 corrects the acquired procedure name to the labeled procedure name. However, this is just one example; the procedure name correction unit 8 may also provide the language model with the action interval information and procedure name information acquired from the procedure name acquisition unit 7, and instruct the language model to check the procedure name indicated by the procedure name information based on the text indicating the purpose of the task and the action label shown in the video. In this case, the language model determines the procedure name for each action segment in a series of steps, based on the text indicating the purpose of the task and the action labels shown in the video. The language model then compares the determined procedure name with the procedure name obtained by the procedure name acquisition unit 7. If the requested procedure name is, for example, "take cup, pour coffee, pour water, stir coffee," and the procedure name obtained by the procedure name acquisition unit 7 is, for example, "take plate, pour coffee, pour water, stir coffee," then the language model will determine that "take cup" and "take plate" are different. The language model then corrects "take plate" to "take cup" in the procedure name obtained by the procedure name acquisition unit 7, and outputs "take cup,pour coffee,pour water,stir coffee" as the corrected procedure name to the presentation processing unit 4.
[0079] Alternatively, the procedure name modification unit 8 may provide the action interval information and procedure name information output from the procedure name acquisition unit 7 to the language model, and obtain the modified procedure name from the language model. In this case, the language model checks whether the flow of action intervals indicated by the action interval information is natural, and based on the result of that check, if the flow of procedure names indicated by the procedure name information is not reasonable, it modifies the procedure names.
[0080] In the video processing device shown in Figure 11, the procedure name acquisition unit 7 provides time-series data to a trained model 9 and obtains the action interval and procedure name from the trained model 9. The procedure name acquisition unit 7 may further provide the video corresponding to each action interval to a VLM (Vision and Language Model) and obtain text information that describes the content of the video corresponding to each action interval. The procedure name acquisition unit 7 can inform the user of the content of the video corresponding to each action section by displaying the text information on a display device (not shown) via the presentation processing unit 4.
[0081] Furthermore, this disclosure allows for free combination of each embodiment, modification of any component in each embodiment, or omission of any component in each embodiment. [Industrial applicability]
[0082] This disclosure includes a procedure group candidate acquisition unit that acquires multiple procedure group candidates, each containing a possible set of steps for performing a task shown in a video and the name of each step in those steps; and an action interval detection unit that detects action intervals in a video, which are sections in the video in which each of the steps for performing a task may be shown, based on the change in the feature quantities of the task shown in the video over time. The video processing device also includes a procedure group candidate selection unit that selects a procedure group candidate corresponding to an action interval detected by the action interval detection unit from among the multiple procedure group candidates acquired by the procedure group candidate acquisition unit, and can acquire the name of the step corresponding to the action interval, making it suitable for video processing devices and video processing methods. [Explanation of symbols]
[0083] 1 Procedure group candidate acquisition unit, 2 Action interval detection unit, 2a Feature extraction unit, 2b Interval group candidate acquisition unit, 3 Procedure group candidate selection unit, 3a Matching processing unit, 3b Procedure group candidate selection processing unit, 4 Presentation processing unit, 5 Procedure group candidate acquisition unit, 6 Feature extraction unit, 7 Procedure name acquisition unit, 8 Procedure name modification unit, 9 Trained model, 21 Procedure group candidate acquisition circuit, 22 Action interval detection circuit, 23 Procedure group candidate selection circuit, 24 Presentation processing circuit, 25 Procedure group candidate acquisition circuit, 26 Feature extraction circuit, 27 Procedure name acquisition circuit, 28 Procedure name modification circuit, 31 Memory, 32 Processor.
Claims
1. A procedure group candidate acquisition unit acquires multiple procedure group candidates, each containing multiple possible steps for performing a task shown in a video, and the name of each step in those steps. An action interval detection unit detects action intervals in the video that are segments in the video in which each of the steps for performing the task may be shown, based on the changes in the feature quantities of the task shown in the video over time. A procedure group candidate selection unit selects a procedure group candidate from among a plurality of procedure group candidates acquired by the procedure group candidate acquisition unit that corresponds to an action interval detected by the action interval detection unit. A video processing device equipped with a video processing device.
2. The aforementioned action interval detection unit, A feature extraction unit extracts the features of the task shown in the aforementioned video and outputs time-series data showing the change in the features over time. The system includes an interval group candidate acquisition unit that acquires multiple interval group candidates, each containing an action interval that may represent a set of steps for performing the task, based on the time-series data output from the feature extraction unit that shows the time changes of the features, The procedure group candidate selection unit, A matching processing unit that compares each procedure group candidate obtained by the procedure group candidate acquisition unit with each interval group candidate obtained by the interval group candidate acquisition unit, Based on the matching results of the matching processing unit, the procedure group candidate selection processing unit selects from among a plurality of interval group candidates acquired by the interval group candidate acquisition unit an interval group candidate that includes an action interval that represents each of the steps in a series of steps for executing the task, and selects from among a plurality of procedure group candidates acquired by the procedure group candidate acquisition unit a procedure group candidate that corresponds to the action interval included in the selected interval group candidate. The image processing apparatus according to claim 1, characterized by its features.
3. The aforementioned action interval detection unit, A feature extraction unit extracts the features of the task shown in the aforementioned video and outputs time-series data showing the change in the features over time. The system includes a trained model that has learned the correspondence between task features and action intervals, and an interval group candidate acquisition unit that, by providing time-series data output from the feature extraction unit, acquires multiple interval group candidates from the trained model, each of which may represent a set of steps for executing the task. The procedure group candidate selection unit, A matching processing unit that compares each procedure group candidate obtained by the procedure group candidate acquisition unit with each interval group candidate obtained by the interval group candidate acquisition unit, Based on the matching results of the matching processing unit, the procedure group candidate selection processing unit selects from among a plurality of interval group candidates acquired by the interval group candidate acquisition unit an interval group candidate that includes an action interval that represents each of the steps in a series of steps for executing the task, and selects from among a plurality of procedure group candidates acquired by the procedure group candidate acquisition unit a procedure group candidate that corresponds to the action interval included in the selected interval group candidate. The image processing apparatus according to claim 1, characterized by its features.
4. The aforementioned matching processing unit, The similarity between each procedure group candidate obtained by the procedure group candidate acquisition unit and each interval group candidate obtained by the interval group candidate acquisition unit is calculated, and the result of the similarity calculation is output to the procedure group candidate selection processing unit as the matching result. The image processing apparatus according to claim 2 or 3, characterized in that it is a video processing apparatus.
5. Presentation Processing Unit presents the action interval included in the interval group candidate selected by the procedure group candidate selection processing unit, and the procedure name of the procedure included in the procedure group candidate selected by the procedure group candidate selection processing unit. The image processing apparatus according to claim 2 or 3, characterized by comprising:
6. The procedure group candidate acquisition unit, The Artificial Intelligence (Generative AI) is given a request to create a series of steps to perform the task shown in the aforementioned video, and the Generative AI obtains the multiple candidate sets of steps. The image processing apparatus according to claim 1, characterized by its features.
7. The procedure group candidate acquisition unit, Based on the text indicating the purpose of the task shown in the video, and the labels indicating the action intervals (sections in which each of the steps for performing the task is shown) and the step names of the steps corresponding to those action intervals, the multiple candidate sets of steps are obtained. The image processing apparatus according to claim 1, characterized by its features.
8. A feature extraction unit extracts features of tasks shown in a video and outputs time-series data showing how these features change over time. A procedure name acquisition unit provides the time-series data output from the feature extraction unit to a trained model and obtains from the trained model action intervals, which are segments in which each of the series of steps for executing the task shown in the video is shown, and the procedure name of the step corresponding to the action interval. An action label is obtained that shows an action section, which is a section in which each of the steps for executing the aforementioned task is shown, and the name of the step corresponding to the action section. Multiple action sections indicated by the action label are compared with multiple action sections obtained by the step name acquisition unit to identify identical action sections. If the name of the step indicated by the action label and the name of the step obtained by the step name acquisition unit do not match within the same action section, the step name obtained by the step name acquisition unit is corrected to the name of the step indicated by the action label. A video processing device equipped with a video processing device.
9. The aforementioned trained model is During training, given time-series data showing the temporal changes in the features of a task shown in a video, and labels indicating the names of the steps corresponding to the action intervals in which each step of the sequence of steps for performing the task is shown, the model learns the correspondence between the features of the task and the names of those steps. During inference, when time-series data output from the feature extraction unit is provided, the system outputs an action interval, which is a section representing each of the steps in a series of procedures for executing the task of extracting features by the feature extraction unit, and the name of the procedure corresponding to that action interval. The image processing apparatus according to claim 8, characterized in that it is a video processing apparatus.
10. The procedure group candidate acquisition unit acquires multiple procedure group candidates, each including multiple possible steps for executing the task shown in the video, and the name of each step in those steps. The action interval detection unit detects action intervals in the video, which are sections in the video that may contain each of the steps for performing the task, based on the changes in the feature quantities of the task shown in the video over time. The procedure group candidate selection unit selects a procedure group candidate from among the multiple procedure group candidates acquired by the procedure group candidate acquisition unit that corresponds to the action interval detected by the action interval detection unit. Image processing methods.
11. The feature extraction unit extracts the features of the tasks shown in the video and outputs time-series data showing the changes in the features over time. The procedure name acquisition unit provides the time series data output from the feature extraction unit to the trained model and obtains from the trained model the action intervals, which are segments in which each of the series of steps for executing the task shown in the video is shown, and the procedure names of the steps corresponding to those action intervals. The procedure name modification unit obtains action labels indicating action sections, which are sections showing each of the steps in a series of steps for executing the task, and the procedure names of the steps corresponding to those action sections. The unit compares the multiple action sections indicated by the action labels with the multiple action sections obtained by the procedure name acquisition unit to identify identical action sections. If the procedure name indicated by the action label and the procedure name obtained by the procedure name acquisition unit do not match within the same action section, the unit modifies the procedure name obtained by the procedure name acquisition unit to the procedure name indicated by the action label. Image processing methods.
Citation Information
Patent Citations
Abnormal transmission program, abnormal transmission method, and information processing device
JP2024032618A
Information processing apparatus and control method of the same
JP2024048120A