Dynamic image brochure creating apparatus, dynamic image brochure creating method, and storage device storing dynamic image brochure creating program

By analyzing the work process manual and dynamic images, the correspondence between objects and actions is generated, which solves the problem that objects and actions in the dynamic image manual are not easy to understand and realizes a clear display of the dynamic image manual.

CN117280339BActive Publication Date: 2026-01-06MITSUBISHI ELECTRIC CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202180098158.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-20
Publication Date
2026-01-06
Estimated Expiration
2041-05-20

AI Technical Summary

Technical Problem

The existing dynamic image manuals have a problem with the difficulty in understanding the correspondence between objects and human movements.

Method used

The document analysis department analyzes the work process document to generate statement information data; the dynamic image analysis department analyzes the dynamic image file to generate object and action information data; the link information generation department extracts the corresponding relationship from the statement and object action information to generate link information data; and the dynamic image manual generation department synthesizes the dynamic image manual data, and the relevant information is displayed on the monitor.

Benefits of technology

It makes the correspondence between objects and human actions clearly visible, improving the comprehensibility of dynamic image manuals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117280339B_ABST
    Figure CN117280339B_ABST
Patent Text Reader

Abstract

The moving image manual creation apparatus (100) has: a document analysis section (101); a moving image analysis section (102); a link information generation section (106) that collects a group of a noun and a verb, i.e., a first group (150), from sentence information data (D101), collects a group of an object and an action, i.e., a second group (160a, 160b), from object information data (D103) and action information data (D104), retrieves the first group and the second group in which a noun and an object correspond and a verb and an action correspond from these groups, and generates link information data (D106) that indicates correspondence between a position (151) within a work flow in which the first group obtained by the retrieval is recorded and a scene (161) within a moving image that contains the second group obtained by the retrieval; and a moving image manual generation section (107) that generates moving image manual data (D107) that causes a display to display a moving image manual that contains a work flow, a moving image, a noun, and a verb, in accordance with the link information data (D106).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an apparatus for creating animated image manuals, a method for creating animated image manuals, and a storage device storing an animated image manual creation program. Background Technology

[0002] A method for creating multimedia manuals, namely animated image manuals, that include usage statements and animated images has been proposed. For example, see Patent Document 1. In this method, the animated image is analyzed to create an event list consisting of time information and object information (e.g., tool name) when the operator uses a tool or component, and the event time information is assigned to the animated image, thereby creating an animated image manual.

[0003] Existing technical documents

[0004] Patent documents

[0005] Patent Document 1: Japanese Patent Application Publication No. 2008-225883 Summary of the Invention

[0006] The problem that the invention aims to solve

[0007] However, in the aforementioned motion picture manual, there is a problem that makes it difficult to understand the correspondence between objects (such as objects, props) and human actions (such as the movement of workers).

[0008] The purpose of this invention is to provide an apparatus, method, and program for creating dynamic image manuals that are easy to understand in relation to the correspondence between objects and human actions.

[0009] Methods for solving problems

[0010] The motion picture manual production apparatus of the present invention is characterized in that it comprises: a document analysis unit that analyzes a work process document containing a work process and generates statement information data representing the structure of statements contained in the work process document; a motion picture analysis unit that analyzes a motion picture file obtained by a camera capturing a person performing a work based on the work process and generates object information data including in-screen coordinate positions and the names of objects contained in the frame images constituting the motion picture, and generates action information data representing the actions of the person, including the coordinate positions and movement directions of the person's body parts contained in the motion picture, wherein the in-screen coordinate positions represent the positions of the object of the work, the props used in the work, and the body parts in each frame image; and a link information generation unit that collects a group of nouns and verbs contained in the statements, i.e., a first group, from the statement information data and generates action information data representing the actions of the person, including the coordinate positions and movement directions of the person's body parts contained in the motion picture, wherein the in-screen coordinate positions represent the positions of the object of the work, the props used in the work, and the body parts in each frame image; and a link information generation unit that collects a group of nouns and verbs contained in the statements, i.e., a first group, from the statement information data and generates action information data representing the actions of the person, including the coordinate positions and movement directions of the person's body parts contained in the motion picture. The system collects the in-screen coordinates, names, and actions of the objects contained in the dynamic image from the body information data and the action information data, i.e., the second group. It retrieves the first group and the second group from the collected first group and the collected second group, where the noun and the name of the object correspond to the verb and the action, and generates link information data indicating the corresponding position in the work process recorded by the retrieved first group and the scene in the dynamic image containing the retrieved second group; and a dynamic image manual generation unit generates dynamic image manual data for displaying dynamic image manual on the display based on the link information data. The dynamic image manual includes the work process, the dynamic image, the nouns near the objects displayed in each frame of the dynamic image, and the verbs near the body parts performing the actions corresponding to the verbs displayed in each frame of the dynamic image.

[0011] The method for creating a dynamic image manual according to the present invention is executed by a dynamic image manual creation device for creating dynamic image manual data. The method comprises the following steps: analyzing a work process document containing a work procedure to generate statement information data representing the structure of statements contained in the work process document; analyzing a dynamic image file obtained by a camera capturing a person performing a task based on the work procedure to generate object information data including the coordinate position within the frame and the names of objects contained in the frames constituting the dynamic image; generating action information data representing the person's actions, including the coordinate position and direction of movement of the person's body parts, contained in the dynamic image, wherein the coordinate position within the frame represents the position of the object of the task, the props used in the task, and the body parts in each frame; and collecting groups of nouns and verbs contained in the statements from the statement information data. Group 1, which collects the in-screen coordinates, names, and actions of the objects contained in the dynamic images from the object information data and the action information data; Group 2, which retrieves the first and second groups from the collected first and second groups, where the nouns and object names correspond and the verbs and actions correspond, and generates corresponding link information data indicating the relationship between the positions in the work process obtained from the first group and the scenes in the dynamic images containing the second group obtained from the second group; and generates dynamic image manual data based on the link information data to display a dynamic image manual on the monitor, the dynamic image manual including the work process, the dynamic images, the nouns near the objects displayed in each frame of the dynamic images, and the verbs near the body parts performing the actions corresponding to the verbs displayed in each frame of the dynamic images.

[0012] Invention Effects

[0013] According to the present invention, it is possible to create dynamic image manuals that are easy to understand in relation to the correspondence between objects and human actions. Attached Figure Description

[0014] Figure 1 This is a functional block diagram that schematically illustrates the structure of the dynamic image manual production apparatus according to Embodiment 1.

[0015] Figure 2 This is a diagram showing an example (one of) of a moving image manual produced by the moving image manual production apparatus of Embodiment 1.

[0016] Figure 3This is a diagram showing an example (second example) of a moving image manual produced by the moving image manual production apparatus of Embodiment 1.

[0017] Figure 4 This is a diagram illustrating an example of the hardware structure of a system (e.g., a computer) that implements the motion picture manual production apparatus and display control device of Embodiment 1.

[0018] Figure 5 This is a diagram illustrating an example of the structure of statement information data generated by the document analysis department.

[0019] Figure 6 This is a diagram illustrating an example of the structure of object information data generated by the object detection unit of the dynamic image analysis unit.

[0020] Figure 7 This is a diagram illustrating an example of the structure of motion information data generated by the motion detection unit of the dynamic image analysis unit.

[0021] Figure 8 This is a diagram illustrating an example of the structure of link information data generated by the link information generation unit.

[0022] Figure 9 This is a flowchart illustrating the generation and processing of statement information data by the document analysis department.

[0023] Figure 10 This is a diagram illustrating an example of a tree structure for statement information data generated by the document analysis department.

[0024] Figure 11 This is a flowchart illustrating the object information data generation process performed by the object detection unit of the dynamic image analysis unit.

[0025] Figure 12 This is a diagram illustrating an example of a tree structure for object information data generated by the object detection unit of the dynamic image analysis unit.

[0026] Figure 13 This is a flowchart illustrating the motion information data generation and processing performed by the motion detection unit of the dynamic image analysis unit.

[0027] Figure 14 This is a diagram illustrating an example of a tree structure for motion information data generated by the motion detection unit of the motion image analysis unit.

[0028] Figure 15 This is a flowchart illustrating the link information data generation process performed by the link information generation unit.

[0029] Figure 16 This is a diagram illustrating the tree structure generation process of the link information data performed by the link information generation unit.

[0030] Figure 17 This is a flowchart illustrating the dynamic image manual generation process performed by the dynamic image manual generation unit.

[0031] Figure 18 This is a flowchart illustrating the display processing of a dynamic image manual performed by a display control device.

[0032] Figure 19 This is a functional block diagram that schematically illustrates the structure of the dynamic image manual production apparatus according to Embodiment 2.

[0033] Figure 20 This is a diagram illustrating an example of the hardware structure of a system (e.g., a computer) that implements the motion picture manual production apparatus and display control device of Embodiment 2.

[0034] Figure 21 This is a flowchart showing the processing performed by the voice analysis unit of the dynamic image manual production apparatus according to Embodiment 2.

[0035] Figure 22 This is a diagram illustrating an example of a tree structure of voice information data generated by the voice analysis unit of the dynamic image manual production apparatus according to Embodiment 2.

[0036] Figure 23 This is a flowchart illustrating the link information data generation process performed by the link information generation unit of the dynamic image manual production apparatus according to Embodiment 2.

[0037] Figure 24 This is a functional block diagram that schematically illustrates the structure of the dynamic image manual production apparatus of Embodiment 3.

[0038] Figure 25 This is a diagram illustrating an example of the hardware structure of a system (e.g., a computer) that implements the motion picture manual production apparatus and display control device of embodiment 3.

[0039] Figure 26 This is a flowchart illustrating the parallel processing of the dynamic image recording unit and the object detection unit in the dynamic image manual production apparatus of Embodiment 3.

[0040] Figure 27 This is a functional block diagram that schematically illustrates the structure of the dynamic image manual production apparatus of Embodiment 4.

[0041] Figure 28 This is a diagram showing the structure of a display control device for displaying a dynamic image manual generated by the dynamic image manual production apparatus of embodiment 4 in AR (augmented reality) glasses.

[0042] Figure 29This is a diagram illustrating an example of the hardware structure of a system (e.g., a computer) that implements the dynamic image manual production apparatus and display control device of embodiment 4.

[0043] Figure 30 This is a flowchart illustrating the processing performed by the overlap position alignment control unit of the dynamic image manual production apparatus according to Embodiment 4. Detailed Implementation

[0044] The following description, with reference to the accompanying drawings, outlines the apparatus, method, and procedure for creating a moving image manual according to various embodiments. These embodiments are merely examples, and the embodiments can be appropriately combined and modified.

[0045] Implementation Method 1

[0046] Figure 1 This is a functional block diagram schematically illustrating the structure of the motion picture manual production apparatus 100 according to Embodiment 1. The motion picture manual production apparatus 100 is an apparatus capable of executing the motion picture manual production method of Embodiment 1. The motion picture manual data D107 produced by the motion picture manual production apparatus 100 is output to the display control device 110. The display control device 110 includes a motion picture reproduction control unit 112 for controlling motion picture reproduction and a motion picture manual display control unit 111 for controlling the display operation of the motion picture manual. The display control device 110 causes the display 120, which is an image display device, to display the motion picture manual. The motion picture manual production apparatus 100, the display control device 110, and the display 120 constitute a motion picture manual prompting system for prompting a person (e.g., an operator) with the motion picture manual. Furthermore, the display control device 110 may also be part of the motion picture manual production apparatus 100.

[0047] like Figure 1 As shown, the motion picture manual generation apparatus 100 includes a document analysis unit 101, a motion picture analysis unit 102, a link information generation unit 106, and a motion picture manual generation unit 107. The motion picture analysis unit 102 includes an object detection unit 103 and a motion detection unit 104.

[0048] The document analysis unit 101 analyzes the work process document file that describes the work process document and generates statement information data D101, which represents the structure of the statements contained in the work process document file. The document analysis unit 101 collects the nouns and verbs contained in the statements. For example, nouns are words that represent the names of objects, and verbs are words that represent the movement of people (e.g., operators).

[0049] The dynamic image analysis unit 102 analyzes dynamic image files that capture dynamic images of a work process. The object detection unit 103 of the dynamic image analysis unit 102 detects objects contained in the dynamic image and generates object information data D103 representing the objects. The motion detection unit 104 of the dynamic image analysis unit 102 detects human motion contained in the dynamic image and generates motion information data D104 representing the motion. The objects include at least one of the following: the object of the work, the props used in the work, and a part of a human body.

[0050] The link information generation unit 106 collects the first group (group of nouns and verbs contained in sentences) from the sentence information data D101, and the second group (group of objects and human actions contained in motion images) from the object information data D103 and the action information data D104. The link information generation unit 106 then collects the first group (described later) from the second group. Figure 5 150 of them) and the second group collected (described later) Figure 6 , Figure 7 In sections 160a and 160b), the first and second groups are retrieved where the noun and object correspond and the verb and action correspond. The link information generation unit 106 generates a link indicating the location (described later) of the work process document containing the first group obtained through the retrieval. Figure 8 151) and the scene contained in the second group of moving images obtained through retrieval (described later) Figure 8 The corresponding link information data D106 between 161) in the middle.

[0051] The motion picture manual generation unit 107 generates motion picture manual data D107 based on the work process document, motion picture file, and link information data D106. This motion picture manual data D107 is used to display a motion picture manual containing work process, motion picture, nouns, and verbs on the display.

[0052] Figure 2 and Figure 3 These figures illustrate examples (one and two) of a moving image manual produced by the moving image manual production apparatus 100 and displayed on the display 120. Figure 2 and Figure 3 In the display, the contents of the work procedure book are displayed on the left half of the monitor 120, and dynamic images are displayed on the right half. Figure 2 The following example illustrates this: Extract the nouns "left hand," "substrate," "right hand," "screwdriver," and "screw" from the statement "Press the substrate with your left hand while rotating the screws at the four corners with the screwdriver in your right hand" in item 3 of the work procedure book. Display the corresponding nouns in the animation near the objects (e.g., substrate, screws), the tools used in the work (e.g., tools), and the body parts of the operator (e.g., left hand, right hand). Figure 3 An example is shown as follows: Verbs, namely "press" and "rotate", are extracted from the sentences in item 3 of the operation process manual, and the corresponding verbs are displayed near the body parts that perform the actions corresponding to the verbs in the moving image.

[0053] Figure 4 It is a diagram showing an example of the hardware structure of a system, namely a computer, that implements the moving image manual production device 100 and the display control device 110. As Figure 4 shown, the computer has a processor for processing information, namely a CPU (Central Processing Unit), 510, a main memory 520 such as a RAM (Random Access Memory), a storage device 530 such as a hard disk drive (HDD) or a solid state drive (SSD), an output interface (I / F) 540, an input I / F 550, a communication I / F 560, and an image processing processor 570. A display 120, a touch panel 121, a keyboard / mouse 122, a camera / microphone 123, a network 125, and a tablet / smartphone 124 are connected to the computer.

[0054] Each function of the moving image manual production device 100 and the display control device 110 can also be implemented by a processing circuit. The processing circuit can be dedicated hardware or a processor that executes programs (such as a moving image manual production program, a display control program, etc.) stored in the main memory.

[0055] Regarding each function of the moving image manual production device 100 and the display control device 110, part of it can be implemented by dedicated hardware, and part of it can be implemented by software or firmware. In this way, the processing circuit can implement Figure 1 the functions of each functional block shown by a combination of hardware, software, firmware, or any one of them.

[0056] Figure 5 It is a diagram showing a structural example of the statement information data D101 generated by the document analysis unit 101. In Figure 5 it, the arrows indicate the definition of the hierarchical structure between data, and the arrow destination points to the lower level. In addition, in Figure 5 it, "procedure" represents one statement. Figure 5 [[ID=​​​This is a diagram illustrating an example of the structure of object information data D103 generated by the object detection unit 103 of the dynamic image analysis unit 102. Figure 6 In the object information data D103, the following example is shown: each "moving image" consists of one or more "frame images", each "frame image" contains one or more "objects (objects or props)", and each "object" consists of one or more coordinates representing the "coordinate position within the screen". In addition, by multiplying the reciprocal of the frame rate value by the frame number, the playback time from the beginning of the moving image can also be represented.

[0058] Figure 7 This is a diagram illustrating an example of the structure of motion information data D104 generated by the motion detection unit 104 of the dynamic image analysis unit 102. Figure 7 In the motion information data D104, the following example is shown: Each "motion image" consists of one or more "frame images", each "frame image" contains one or more "objects (objects or props)", and each "object" consists of one or more coordinates representing the "coordinate position in the screen", the root coordinate of the arrow representing the "movement direction" (i.e., the coordinate of the rear end of the arrow) and the destination coordinate of the arrow (i.e., the coordinate of the front end of the arrow).

[0059] Figure 8 This is a diagram showing an example of the structure of the link information data D106 generated by the link information generation unit 106. Figure 8 The link information data D106 is from Figure 5 Statement information data D101, Figure 6 Object information data D103 and Figure 7 The motion information data D104 constitutes the data. Figure 8 The link information data D106 shows examples of each "procedure" consisting of an "object (object)", an "object (prop or body part)", and an "action".

[0060] Figure 9 This is a flowchart illustrating the generation process of statement information data D101 performed by the document analysis unit 101. First, the document analysis unit 101 reads the work flow document (step S101), and determines whether the chapter number (i.e., chapter number and section number) of the read statement is the last chapter number of the work flow document (step S102). If it is the last chapter number (step S102: Yes), the document analysis unit 101 ends the generation process of statement information data D101. If it is not the last chapter number (step S102: No), the document analysis unit 101 determines the generation process based on the hierarchical tree of chapter and section numbers (described later). Figure 10 As shown, generate nodes (step S103).

[0061] The document analysis unit 101 determines whether it is the final "procedure" (i.e., statement or text) (step S104). If it is the final "procedure" (step S104: Yes), the processing returns to step S102. If it is not the final "procedure" (step S104: No), the document analysis unit 101 cuts out the "procedure" with the section number (step S105).

[0062] The document analysis unit 101 extracts words through morpheme analysis and determines the word class (step S106). The document analysis unit 101 performs grammatical analysis using a dictionary, generating nodes according to each object (noun), prop (noun), and action (verb) (step S107). The document analysis unit 101 repeats steps S104 to S107 until the final "procedure".

[0063] Figure 10 This is a diagram illustrating an example of the tree structure of the statement information data D101 generated by the document analysis unit 101. (See diagram for example.) Figure 10 As shown, a "document" consists of one or more nodes called "chapter", each "chapter" consists of one or more nodes called "section", each "section" consists of one or more nodes called "procedure" (statement or text), and each "procedure" consists of one or more nodes called object and one or more nodes called action.

[0064] Figure 11 This is a flowchart illustrating the process of generating object information data D103 by the object detection unit 103 of the dynamic image analysis unit 102. First, the object detection unit 103 reads a dynamic image file (i.e., an image file) captured by a camera (step S111), and determines whether the frame image of the read dynamic image file is the frame number of the last frame image (step S112). If it is the last frame number (step S112: Yes), the object detection unit 103 ends the process of generating object information data D103. If it is not the last frame number (step S112: No), the object detection unit 103 generates nodes according to each frame image (described later). Figure 12 (as shown in step S113).

[0065] Next, the object detection unit 103 detects objects (e.g., objects, props, body parts) in the image based on image analysis processing. The object detection unit 103 determines whether there are any undetected objects, i.e., other objects (step S115). If no other objects exist (step S115: No), the processing returns to step S112. If other objects exist (step S115: Yes), the object detection unit 103 generates nodes for each object (described later). Figure 12 As shown (step S116), generate coordinate position nodes for each object (described later). Figure 12(as shown in step S117).

[0066] Figure 12 This is a diagram illustrating an example of the tree structure of the object information data D103 generated by the object detection unit 103. (See diagram for example.) Figure 12 As shown, a "dynamic image" is composed of one or more nodes, namely "frame images" (frame numbers 0001 to 1234), each "frame image" is composed of one or more nodes, namely "objects", and each "object" is composed of one or more nodes, namely "coordinate positions".

[0067] Figure 13 This is a flowchart illustrating the motion information data D104 generation process performed by the motion detection unit 104. First, the motion detection unit 104 reads in a motion image file (step S121) and determines whether the frame image of the read motion image file is the frame number of the last frame image (step S122). If it is the last frame number (step S122: Yes), the motion detection unit 104 ends the motion information data D104 generation process. If it is not the last frame number (step S122: No), the motion detection unit 104 generates nodes according to each frame image (described later). Figure 14 (as shown in step S123).

[0068] Next, the motion detection unit 104 detects motion (e.g., hand movement) in the image based on the skeleton extraction process (step S124). Then, the motion detection unit 104 generates each motion node (described later) according to the motion generation node. Figure 14 (as shown in step S125), nodes are generated according to the coordinate position and movement direction of each action (described later). Figure 14 (as shown in step S126).

[0069] Figure 14 This is a diagram illustrating an example of the tree structure of motion information data D104 generated by the motion detection unit 104. (See diagram for example.) Figure 14 As shown, a "dynamic image" is composed of one or more nodes, namely "frame images" (frame numbers 0001 to 1234), each "frame image" is composed of one or more nodes, namely "action" (including posture), and each "action" is composed of one or more nodes, namely "coordinate position" and "movement direction".

[0070] Figure 15 This is a flowchart illustrating the link information data D106 generation process performed by the link information generation unit 106. First, the link information generation unit 106 reads the tree structure of the statement information data D101 (e.g., ...). Figure 10(Step S131) ​​Determine whether the tree structure of the read statement information data D101 is the tree structure of the last chapter number (i.e., chapter number and section number) (Step S132). If it is the last chapter number (Step S132: Yes), the link information generation unit 106 saves the link information data D106 to the storage device 530 (Step S133) and ends the generation process of the link information data D106. If it is not the last chapter number (Step S132: No), the link information generation unit 106 obtains the node of the group of the three elements of the "procedure" {object as object, prop as object, and human action} (Step S134).

[0071] Next, the link information generation unit 106 retrieves the tree (e.g., from the object information data D103) from the object information data D103. Figure 12 ) and the tree of motion information data D104 (e.g. Figure 14 The link information generation unit 106 generates a hybrid tree of object information and motion information data (step S135), and determines whether a scene with consistent three elements has started (step S136). If a scene with consistent three elements has started (step S136: Yes), the link information generation unit 106 saves the scene information of the starting scene (step S137) and returns the process to step S136. If no scene with consistent three elements exists (step S136: No), the link information generation unit 106 determines whether there is an ending scene with consistent three elements (step S138).

[0072] In determining whether there is an ending scene with three identical elements (step S138), if an ending scene exists (step S138: Yes), the link information generation unit 106 saves the scene information of the ending frame (step S139).

[0073] Next, the link information generation unit 106 generates a link between the "procedure" node and the scene information node (the start time and end time of the scene) (step S140). The link information generation unit 106 generates links for the coordinate positions and movement directions of the nodes for the three elements of the "procedure" (step S141).

[0074] Figure 16 This diagram illustrates the tree structure generation process of the link information data D106 performed by the link information generation unit 106. Figure 16 The link information generation unit 106 creates a tree of link information data D106 by linking the frame image of the "procedure" of the tree of statement information data D101 and the mixed tree of object information / action information data.

[0075] Figure 17This is a flowchart illustrating the motion picture manual generation process performed by the motion picture manual generation unit 107. The motion picture manual generation unit 107 reads the work process document file (step S151), reads the link information data D106 (step S152), and determines whether the chapter number of the read work process document file is the last chapter number (step S153). If it is the last chapter number (step S153: Yes), the motion picture manual generation unit 107 ends the generation process of the motion picture manual data D107. If it is not the last chapter number (step S153: No), the motion picture manual generation unit 107 determines the text position of that chapter number within the work process document (step S154).

[0076] Next, the motion picture manual generation unit 107 obtains the scene information (e.g., the start time of the playback) corresponding to the chapter number from the link information data, and generates a link (e.g., an embedded link code) for the scene information at the text position of the chapter number.

[0077] Figure 18 This is a flowchart illustrating the display processing of the motion picture manual performed by the display control device 110. First, the display control device 110 accepts the designation of a chapter number (e.g., a user click) on the workflow screen of the motion picture manual data D107 (step S161). Next, the display control device 110 jumps to the playback start position on the job motion picture screen by executing the link code (step S162).

[0078] Next, the display control device 110 reads one frame of image from the dynamic image (step S163). The display control device 110 determines whether the playback position is the playback end position (step S164). If it is the playback end position (step S164: Yes), the playback of the dynamic image is stopped (step S169).

[0079] If the reproduction position is not the reproduction end position (step S164: No), the control device 110 will accept the designation of the object, i.e. the target or prop and action in the "procedure" on the work process book screen (click) (step S165).

[0080] Next, the display control device 110 refers to the link information table to obtain the coordinate position and movement direction information of the specified item (step S166). Next, the display control device 110 overlaps the desired position with the emphasis mark in the current image frame (step S167). Next, the display control device 110 overlaps the desired position with the emphasis mark in the current image frame and reproduces the dynamic image (step S168).

[0081] As explained above, in the correspondence between the "procedure" section of the work process book and the corresponding scene of the work animation (e.g., the correspondence between time points), the motion picture manual production apparatus 100 of Embodiment 1 searches for and compares data that are consistent with at least two types of information: the object as an object, the prop as an object, and the human action. Therefore, the error rate of comparison can be reduced.

[0082] Furthermore, in the dynamic image manual production device 100 of Embodiment 1, the "procedures" in the "chapter" of the work process book and the corresponding scenes on the dynamic image are linked by the link information data D106, which can be uniquely determined. Therefore, when the operator specifies any part of each "procedure" (one sentence) in the "chapter" of the work process book or any part of the object (~o), tool (~de), or action (~suru) written in the "procedure" by means of clicking the mouse, it can be highlighted immediately and individually (framed, colored, blinked, overlapping arrows, etc.). The highlighted display indicates the object (material, part, etc.) specified as the object, the tool (tool, right hand, etc.) specified as the object, and the action (direction, degree) specified as the object on the image of the corresponding scene of the work dynamic image.

[0083] Furthermore, in the dynamic image manual production apparatus 100 of Embodiment 1, the "procedure" in the "chapter" of the work process book and the corresponding scene on the dynamic image of the work process are bidirectionally linked by "link information data" so that they can be uniquely identified. Therefore, the time point of the scene is determined by temporarily stopping the reproduction on the dynamic image of the work process, and the scene is identified as the "procedure" part in the "chapter" of the work process book, so the screen can be automatically changed and highlighted.

[0084] Furthermore, in Embodiment 1, the object detection unit 103 and the motion detection unit 104 detect and maintain object information and worker motion information within each image based on the timeline of the work motion image within the object information data D103 and motion information data D104. Therefore, by using character or voice input representing keywords related to the work content—objects, props, and human actions—the desired work scene can be retrieved. Furthermore, as a result, jumps and starting point reproduction of the desired scene within the motion image can be performed easily and accurately.

[0085] Implementation Method 2

[0086] Figure 19 This is a functional block diagram that schematically illustrates the structure of the dynamic image manual production apparatus 200 according to Embodiment 2. Figure 19 In the middle, to and Figure 1 The structures shown are the same or the corresponding structural labels are the same as those shown. Figure 1The same reference numerals are shown. The motion picture manual production apparatus 200 is an apparatus capable of executing the motion picture manual production method of Embodiment 2. The motion picture manual production apparatus 200 differs from the motion picture manual production apparatus 100 of Embodiment 1 in that the motion picture analysis unit 202 has a voice analysis unit 105 that analyzes the voice of the motion picture file, and the link information generation unit 106 also uses voice information data D105. The motion picture manual data D107 produced by the motion picture manual production apparatus 200 is output to the display control device 110. The motion picture manual production apparatus 200, the display control device 110, and the display 120 constitute a motion picture manual prompting system for prompting a person (e.g., an operator) with a motion picture manual. Furthermore, the display control device 110 may also be part of the motion picture manual production apparatus 200.

[0087] Figure 20 This is a diagram illustrating an example of the hardware structure of a system (e.g., a computer) 200a that implements the dynamic image manual production apparatus 200 and the display control device 110. Figure 20 In the middle, to and Figure 4 The structures shown are the same or the corresponding structural labels are the same as those shown. Figure 4 The labels shown are the same. Figure 20 System 200a and Figure 4 The difference with System 100a is that it analyzes the voice of the moving image file and uses the voice information data D105 to create a moving image manual.

[0088] Figure 21 This is a flowchart illustrating the processing performed by the speech analysis unit 105 of the motion picture manual production apparatus 200. First, the speech analysis unit 105 reads the speech from the motion picture file (step S201), and determines whether the frame image of the read motion picture file is the last frame image (step S202). If it is the last frame image (step S202: Yes), the speech analysis unit 105 ends the generation process of the speech information data D105. If it is not the last frame number (step S202: No), the speech analysis unit 105 obtains the speech start time (step S203), and generates nodes according to each frame image (described later). Figure 22 (as shown in step S204).

[0089] Next, the speech analysis unit 105 converts the speech into text based on speech recognition processing (step S205). Then, the speech analysis unit 105 determines the object (~を), the prop (~で), and the action (~する) for each sentence of "procedure" (step S206), and generates nodes (described later) for each object, prop, and action. Figure 22 (as shown in step S207).

[0090] Figure 22 This is a diagram illustrating an example of the tree structure of the speech information data D105 generated by the speech analysis unit 105. (See diagram for example.) Figure 22 As shown, a "moving image" consists of one or more nodes, namely "frame images" (e.g., frame numbers 0001 to 1234), and each "frame image" consists of one or more nodes, namely "object (object object)", "object (prop)", and "action".

[0091] Figure 23 This is a flowchart illustrating the link information data D106 generation process performed by the link information generation unit 106 of the motion picture manual production apparatus 200. First, the link information generation unit 106 reads the tree structure of the statement information data D101 (step S211), and determines whether the tree structure of the read statement information data D101 is the tree structure of the last chapter number (i.e., chapter number and section number) (step S212). If it is the last chapter number (step S212: Yes), the link information generation unit 106 saves the link information data D106 to the storage device 530 (step S213), ending the link information data D106 generation process. If it is not the last chapter number (step S212: No), the link information generation unit 106 obtains the node of the group of three elements of the "procedure" {object as an object, prop as an object, and human action} (step S214).

[0092] Next, the link information generation unit 106 retrieves the tree from the object information data D103 and the tree from the motion information data D104 (e.g., Figure 14 The link information generation unit 106 generates a mixed tree of object information / action information / voice information data composed of object information and voice information data (step S215), and determines whether a scene with consistent three elements has started (step S216). If a scene with consistent three elements has started (step S216: Yes), the link information generation unit 106 saves the scene information of the starting scene (step S217) and returns the processing to step S216. If no scene with consistent three elements exists (step S216: No), the link information generation unit 106 determines whether there is an ending scene with consistent three elements (step S218).

[0093] In determining whether there is an ending scene with three identical elements (step S218), if an ending scene exists (step S218: Yes), the link information generation unit 106 saves the scene information of the ending frame (step S219).

[0094] Next, the link information generation unit 106 generates a link between the "procedure" node and the scene information node (the start time and end time of the scene) (step S220). The link information generation unit 106 generates links for the coordinate positions and movement directions of the nodes for the three elements of the "procedure" (step S221).

[0095] As explained above, the motion picture manual production apparatus 200 of Embodiment 2 is equipped with a speech analysis unit 105. The speech analysis unit 105 analyzes the speech within the motion picture, extracts the speech within the motion picture (e.g., speech related to objects, props, and human actions), and outputs structured speech information data D105 along the timeline of the motion picture. Therefore, for example, if it is a cooking motion picture, when the operator performs the operation by explaining the steps of the operation in the motion picture, a cooking motion picture manual with voice explanation can be produced.

[0096] Furthermore, the link information generation unit 106 uses voice information data D105 to perform the correspondence between the work process book and the work animation image, thereby improving the accuracy of the correspondence processing between the work process book and the animation image.

[0097] For matters other than those described above, Implementation 2 is the same as Implementation 1.

[0098] Implementation Method 3

[0099] Figure 24 This is a functional block diagram that schematically illustrates the structure of the dynamic image manual production apparatus 300 according to Embodiment 3. Figure 24 In the middle, to and Figure 19 The structures shown are the same or the corresponding structural labels are the same as those shown. Figure 19 The same reference numerals are shown. The motion picture manual production apparatus 300 is an apparatus capable of performing the motion picture manual production method of Embodiment 3. The motion picture manual production apparatus 300 differs from the motion picture manual production apparatus 200 of Embodiment 2 in that it has a motion picture recording unit 308 for recording motion pictures captured by a camera, and the recorded motion pictures are analyzed by the motion picture analysis unit 202. The motion picture manual data D107 produced by the motion picture manual production apparatus 300 is output to the display control device 110. The motion picture manual production apparatus 300, the display control device 110, and the display 120 constitute a motion picture manual prompting system for prompting a person (e.g., an operator) with a motion picture manual. In addition, the display control device 110 may also be part of the motion picture manual production apparatus 300.

[0100] Figure 25This is a diagram illustrating an example of the hardware structure of a system (e.g., a computer) 300a that implements the dynamic image manual production apparatus 300 and the display control device 110 of Embodiment 3. Figure 25 In the middle, to and Figure 20 The structures shown are the same or the corresponding structural labels are the same as those shown. Figure 20 The labels shown are the same. Figure 25 System 300a and Figure 20 The difference in System 200a is that it uses the motion image file received from the motion image recording unit 308 by the motion image analysis unit 202 to create motion image manual data.

[0101] Figure 26 This is a flowchart illustrating the parallel processing of the dynamic image recording unit 308 and the object detection unit 103 in the dynamic image manual production apparatus 300 of Embodiment 3. Figure 26 In the middle, to and Figure 11 The processing shown is the same as the processing label. Figure 11 The step numbers are shown. The difference between the motion picture manual production apparatus 300 and the motion picture manual production apparatus 200 of Embodiment 2 is that the motion picture recording unit 308 reads the images captured by the camera and writes the motion picture file into the storage device 530 until the last frame, and the motion picture analysis unit 202 reads the motion picture file received from the motion picture recording unit 308.

[0102] The motion image manual production apparatus 300 of embodiment 3 has a motion image recording program, in which motion images captured by a camera are recorded in a storage device 530. Then, the motion image manual production apparatus 300 performs a mapping between the work process manual and the currently captured motion images of the work, according to the motion image manual generation program. Therefore, the operator can start / stop camera recording at the work site and check the mapping between their own motion image and the work content in the work process manual, as well as improvement points, on the motion image manual (in this case, the motion image part is the operator's own motion image) produced and displayed on site.

[0103] Furthermore, the dynamic image manual production device 300 newly adds a camera to the work site to capture the operator's work status and compares the work procedure manual with the captured dynamic images. Therefore, it can warn of omissions and errors in the work items based on the content of the work procedure manual and correct the errors. In this way, the dynamic image manual production device 300, by having the function of recording dynamic images, is not limited to the prompts in the dynamic image manual, but can also play an educational role for the operator.

[0104] For matters other than those described above, Implementation 3 is the same as Implementation 1 or 2 described above.

[0105] Implementation Method 4

[0106] Figure 27 This is a functional block diagram that schematically illustrates the structure of the dynamic image manual production apparatus 400 according to Embodiment 4. Figure 27 In the middle, to and Figure 24 The structures shown are the same or the corresponding structural labels are the same as those shown. Figure 24 The same reference numerals are shown. The motion picture manual production apparatus 400 is an apparatus capable of executing the motion picture manual production method of Embodiment 4. The motion picture manual production apparatus 400, the display control device 410, and the AR (augmented reality) glasses 420 constitute a motion picture manual prompting system for prompting a person (e.g., a worker) with a motion picture manual. The difference between the motion picture manual prompting system of Embodiment 4 and the motion picture manual prompting system of Embodiment 3 is that the AR glasses 420 and the display control device 410 that displays images in the AR glasses 420 are used. Furthermore, the display control device 410 may also be part of the motion picture manual production apparatus 400.

[0107] Figure 28 This diagram illustrates the structure of a display control device 410 for displaying a motion picture manual generated by the motion picture manual production apparatus 400 of Embodiment 4 in AR glasses 420. AR glasses 420, also known as smart glasses, have the function of allowing a person to simultaneously see the real world in front of them and an image superimposed on the real world (e.g., explanatory text superimposed on real-world objects). Furthermore, the AR glasses have a camera 421 that captures motion pictures in the same direction as the gaze of the person wearing the AR glasses. In Embodiment 4, the display control device 410 displays an AR image by superimposing the motion picture manual or a portion thereof on an object in the real world seen at the destination of the person's gaze. The AR image includes, for example, display elements such as boxes and arrows that emphasize real-world objects. The display control device 410 controls the display states, such as the color of the display elements and the presence or absence of flashing of the display elements. To achieve this function, the display control device 410 has a position alignment section that aligns the CG (Computer Graphics) with the position of the real scene, and an overlap section that displays the CG superimposed on the camera image or the real world. The position alignment section and the overlapping section constitute the overlapping position alignment control section 113. The position alignment process is performed, for example, according to the overlapping position alignment procedure. In the position alignment process, the following position alignment process for emphasis display (overlap display) is performed: the dynamic image reflected in the camera 421 is analyzed, and the position information seen from the operator's line of sight is calculated successively based on the position information of each object reflected in the camera image.

[0108] Figure 29 This is a diagram illustrating an example of the hardware structure of a system (e.g., a computer) 400a that implements the dynamic image manual production apparatus 400 and display control device 410 of Embodiment 4. Figure 29 In the middle, to and Figure 25 The structures shown are the same or the corresponding structural labels are the same as those shown. Figure 25 The labels shown are the same. Figure 29 System 400a and Figure 25 The difference with System 300a is that it displays a dynamic image manual in AR glasses 420.

[0109] Figure 30 This is a flowchart illustrating the process performed by the overlap position alignment control unit 113 of the motion picture manual production apparatus 400 according to Embodiment 4. First, the motion picture recording unit 308 reads the frame image for the motion picture manual captured by the camera (step S401). If an end indication is present (step S402: yes), the overlap position alignment control unit 113 causes the AR glasses 420 to display the perspective image, ending the AR image display process.

[0110] In the overlapping position alignment control unit 113, if there is no end indication (step S402: no), the object detection unit 103 detects objects (e.g., objects / props) in the frame image (step S404) and obtains the position information of each detected object in the frame image (step S405).

[0111] Next, the overlap position alignment control unit 113 controls the position alignment of the posture correction (step S406), overlaps (i.e. synthesizes) at appropriate positions of each detected object as CG of the AR image (step S407), and displays it on the screen of the AR glasses 420 (step S408).

[0112] As described above, the device for creating a dynamic image manual according to Embodiment 4 includes AR glasses 420, a camera 421 that captures the work site from the operator's perspective, and AR glasses 420 that have a transparent (see-through) screen. Furthermore, the operator can see (e.g., through the transparent screen) the real world in front of them and can overlay digital (CG) data such as text, images, or dynamic images onto the real world.

[0113] For example, beginners can hear the "procedures" (task content) of the task manual through the speaker of the camera 421 attached to the AR glasses 420, and the real-world objects (such as objects or props) seen at the destination of their vision are superimposed on the display components for emphasis in a form that matches the display position. The display components include, for example, a frame surrounding the area to be emphasized, the color of the area to be emphasized, the blinking of the frame and other display components, and arrows indicating the area to be emphasized.

[0114] By using AR glasses 420, objects (props, etc.) seen at the operator's line of sight are highlighted and displayed in a way that matches their position. Therefore, even beginners can easily and accurately identify parts, materials, props, etc. visually. This helps avoid errors and reduces confusion, thus enabling more efficient work.

[0115] Regarding matters other than those mentioned above, Implementation 4 is the same as any one of Implementations 1 to 3.

[0116] Variations

[0117] In the above implementation, an example of data being a tree structure was described; however, data can also be a structure other than a tree structure.

[0118] Label Explanation

[0119] 100, 200, 300, 400: Dynamic image manual production device; 101: Document analysis unit; 102, 202: Dynamic image analysis unit; 103: Object detection unit; 104: Motion detection unit; 105: Voice analysis unit; 106: Link information generation unit; 107: Dynamic image manual generation unit; 110, 410: Display control device; 120: Display; 150: Group 1; 160a, 160b: Group 2; 420: AR glasses; 421: Camera; 510: CPU; 520: Main memory; 530: Storage device.

Claims

1. A moving image album creating apparatus characterized by comprising: The dynamic image manual creation device has: a document analysis section that analyzes a job flow document in which a job flow is described, and generates sentence information data indicating the structure of a sentence included in the job flow document; a dynamic image analysis section that analyzes a dynamic image file of a dynamic image obtained by a person performing a job based on the job flow with a camera, generates object information data including an in-screen coordinate position and the name of an object included in a frame image constituting the dynamic image, and generates action information data indicating the action of the person including the coordinate position of a body part of the person and the moving direction in the dynamic image, wherein the in-screen coordinate position indicates the position of an object of the job as the object, a prop used in the job, and the body part in each frame image; a link information generation section that collects a group of a noun and a verb included in the sentence, i.e., a first group, from the sentence information data, collects the in-screen coordinate position of the object and the name of the object and a group of the action, i.e., a second group, from the object information data and the action information data, retrieves the first group and the second group in which the noun and the name of the object correspond and the verb and the action correspond from the collected first group and the collected second group, and generates link information data indicating the correspondence between the position in the job flow in which the first group obtained by the retrieval is described and the scene in the dynamic image in which the second group obtained by the retrieval is included; and a dynamic image manual generation section that generates dynamic image manual data that causes a display to display a dynamic image manual including the job flow, the dynamic image, the noun displayed near the object in each frame image of the dynamic image, and the verb displayed near the body part performing the action corresponding to the verb in each frame image of the dynamic image, based on the link information data.

2. The dynamic image manual creation device according to claim 1, wherein the noun is a word indicating the name of the object, the verb is a word indicating the movement of the person.

3. The dynamic image manual creation device according to claim 1 or 2, wherein the dynamic image analysis section includes a voice analysis section that generates voice information data indicating the voice included in the dynamic image file, the link information generation section collects a voice keyword included in the voice information data, the link information generation section retrieves the noun and the verb corresponding to the voice keyword from the sentence information data, and generates the link information data indicating the correspondence between the position of the noun and the verb in the job flow, the scene in the dynamic image, and the voice keyword obtained by the retrieval.

4. The dynamic image manual creation device according to any one of claims 1 to 3, wherein The moving image manual creation device further has a moving image recording section that records, in a storage device, a shot file obtained by a camera shooting the person, The moving image file is the shot file recorded in the storage device.

5. The moving image manual creation device according to any one of claims 1 to 3, wherein The moving image manual creation device further has a display control device, The display control device causes a display to display an augmented reality image in which the moving image manual or a part of the moving image manual is overlaid with a real object seen at a destination of a line of sight of the person as augmented reality information.

6. The moving image manual creation device according to claim 5, wherein The augmented reality image contains a display section that emphasizes the real object, The display control device switches a display state of the display section.

7. A dynamic image album creating method, which is executed by a dynamic image album creating apparatus that creates dynamic image album data, characterized by comprising: The moving image manual creation method has the following steps: analyzing a job flow book file that records a job flow, and generating sentence information data that indicates a structure of a sentence included in the job flow book file; analyzing a moving image file of a moving image obtained by a camera shooting a person who performs a job based on the job flow, and generating object information data that includes an in-screen coordinate position and a name of an object included in a frame image that constitutes the moving image, and generating action information data that indicates a motion of the person included in the moving image, including a coordinate position of a body part of the person and a moving direction; collecting a group of a noun and a verb included in the sentence from the sentence information data as a first group, collecting the in-screen coordinate position of the object and the name from the object information data and the action information data as a second group, retrieving the first group and the second group in which the noun and the name of the object correspond and the verb and the action correspond from the collected first group and the collected second group, and generating link information data that indicates a correspondence between a position within the job flow that records the first group obtained by the retrieval and a scene within the moving image that includes the second group obtained by the retrieval; and generating, from the link information data, moving image manual data that causes a display to display a moving image manual that includes the job flow, the moving image, the noun displayed near the object in each frame image of the moving image, and the verb displayed near the body part that performs the action corresponding to the verb in each frame image of the moving image.

8. A storage device storing a moving image album creating program, characterized by comprising: The moving image manual creation program causes a computer to execute the following steps: analyzing a job flow book file that records a job flow, and generating sentence information data that indicates a structure of a sentence included in the job flow book file; analyzing a moving image file of a moving image obtained by a camera shooting a person who performs a job based on the job flow, and generating object information data that includes an in-screen coordinate position and a name of an object included in a frame image that constitutes the moving image, and generating action information data that indicates a motion of the person included in the moving image, including a coordinate position of a body part of the person and a moving direction; The dynamic image file of the dynamic image taken by the person performing the work based on the work flow is analyzed to generate object information data including in-frame coordinate positions and names of objects included in frame images constituting the dynamic image, and to generate action information data indicating the person's actions including coordinate positions of the person's body parts and moving directions included in the dynamic image, wherein the in-frame coordinate positions indicate positions of objects of the work, props used in the work, and the body parts in each frame image as the objects; A group of nouns and verbs included in the sentence, i.e., a first group, is collected from the sentence information data, a group of the in-frame coordinate positions of the objects and the names and the actions included in the dynamic image, i.e., a second group, is collected from the object information data and the action information data, the first group and the second group in which the nouns correspond to the names of the objects and the verbs correspond to the actions are searched for from the collected first group and the collected second group, and link information data indicating correspondence between positions in the work flow in which the searched first group is recorded and scenes in the dynamic image in which the searched second group is included is generated; and Dynamic image manual data is generated from the link information data to cause a display to display a dynamic image manual including the work flow, the dynamic image, the nouns displayed near the objects in each frame image of the dynamic image, and the verbs displayed near the body parts performing the actions corresponding to the verbs in each frame image of the dynamic image.

Citation Information

Patent Citations

  • Data processor and program for the same

    JP2008225883A

  • Program for creating work assistance data

    CN105745586A

  • Video image processing device, video image processing system, and video image processing method

    CN111279388A