Computer system and program for supporting creation of electronic manual
The computer system and program streamline the creation of electronic manuals by converting video content into structured text and images, addressing inefficiencies in manual creation and reducing the time and effort needed.
Patent Information
- Application Number
- PCT/JP2024/029406
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-20
- Filing Date
- 2024-08-20
- Publication Date
- 2025-08-28
AI Technical Summary
Creating electronic manuals, especially those that include video content, requires significant time and effort, making the process inefficient.
A computer system and program that assist in the creation of electronic manuals by receiving videos, converting them into structured text, dividing them into sub-videos or still images, and generating the manual based on these components, with user input for editing and finalization.
Reduces the time and effort required to create electronic manuals by automating the conversion and structuring process, allowing for efficient generation and customization.
Smart Images

Figure JP2024029406_28082025_PF_FP_ABST
Abstract
Description
Computer system and program for supporting the creation of electronic manuals
[0001] The present invention relates to a computer system and a program for supporting the creation of an electronic manual.
[0002] 2. Description of the Related Art It has been known for some time that electronic manuals are created and used for the purpose of improving work efficiency (see, for example, Patent Document 1).
[0003] WO 2017 / 183064
[0004] However, creating an electronic manual still requires time and effort, and creating an electronic manual that includes video in particular requires a considerable amount of time and effort.
[0005] The present invention has been made in consideration of the above-mentioned problems, and aims to reduce the time and effort required to create an electronic manual by providing a computer system and program for assisting in the creation of an electronic manual.
[0006] In one aspect of the present invention, a computer system of the present invention is a computer system for assisting in the creation of an electronic manual, and the computer system comprises: means for receiving one or more videos; means for receiving information indicating conditions for converting the videos into a plurality of steps; means for generating structured text for constituting the plurality of steps from audio contained in the one or more videos based on the conditions, the structured text including at least a title or description of each of the plurality of steps; means for dividing the one or more videos into a plurality of sub-videos or still images based at least on the one or more videos and the structured text; and means for provisionally generating the electronic manual based on the structured text and the plurality of sub-videos or still images.
[0007] In one embodiment of the present invention, the conditions may include a limit on the number of steps.
[0008] In one embodiment of the present invention, the conditions may further include a limit on the number of characters in the title and / or a limit on the number of characters in the description.
[0009] In one embodiment of the present invention, the audio included in the one or more videos may be audio that indicates the steps of the electronic manual.
[0010] In one embodiment of the present invention, the provisionally generated electronic manual may not include audio included in the one or more videos.
[0011] In one embodiment of the present invention, dividing the one or more videos into a plurality of sub-videos or still images based at least on the one or more videos and the structured text may include dividing the one or more videos into a plurality of candidate sub-videos based at least on the one or more videos and the structured text, and converting a candidate sub-video among the plurality of candidate sub-videos that has audio exceeding a predetermined volume for a predetermined period of time but does not show any change in image into a still image based on the candidate sub-video.
[0012] In one embodiment of the present invention, dividing the one or more videos into a plurality of sub-videos or still images based at least on the one or more videos and the structured text may include identifying timing of scene changes based on the structured text, and generating the plurality of sub-videos or still images by dividing the one or more videos based on the timing of the scene changes.
[0013] In one embodiment of the present invention, identifying the timing of the scene change based on the structured text may include identifying a break in the content of the structured text based on the structured text, and identifying a timing in the audio corresponding to the break in the structured text as the timing of the scene change.
[0014] In one embodiment of the present invention, dividing the one or more videos into multiple sub-videos or still images based at least on the one or more videos and the structured text may further include identifying timings of large image changes in the one or more videos, identifying timings of audio breaks, and generating the multiple sub-videos or still images by dividing the one or more videos at timings where the timings of large image changes, the timings of the scene changes, and the timings of the audio breaks coincide.
[0015] In one embodiment of the present invention, generating the structured text from audio contained in the one or more videos based on the conditions may include converting the audio into text by transcribing the audio contained in the one or more videos, and generating the structured text based on the text converted from the audio and the conditions.
[0016] In one embodiment of the present invention, the computer system may further include means for receiving a first user input indicating a desire to edit the provisionally generated electronic manual; means for identifying, in response to receiving the first user input, a candidate division time period between steps of the provisionally generated electronic manual, wherein within the candidate division time period, a user can adjust the division position between steps of the provisionally generated electronic manual; means for presenting the candidate division time period; means for receiving a second user input for adjusting the division position between steps of the provisionally generated electronic manual within the candidate division time period; and means for editing the provisionally generated electronic manual based on the second user input.
[0017] In one embodiment of the present invention, identifying the time periods of the segmentation candidates may include identifying the time periods of the segmentation candidates based on the structured text and audio contained in the one or more videos.
[0018] In one embodiment of the present invention, identifying the time period of the division candidate based on the structured text and the audio contained in the one or more videos may include identifying the playback time of the audio corresponding to each step of the plurality of steps based on the structured text, and identifying the time period of the division candidate based on the playback time of the audio corresponding to each step.
[0019] In one embodiment of the present invention, the computer system may further include means for receiving a third user input for executing book generation of the electronic manual, and means for executing book generation of the electronic manual in response to receiving the third user input.
[0020] In one embodiment of the present invention, the computer system may further include means for determining whether the one or more videos contain audio, and means for warning a user that the one or more videos do not contain audio if it is determined that the one or more videos do not contain audio.
[0021] In one embodiment of the present invention, the audio included in the one or more videos may be in a colloquial style, and the title and description may be in a written style.
[0022] In one embodiment of the present invention, the computer system may further comprise means for generating audio data for reading the structured text aloud.
[0023] In one embodiment of the present invention, the computer system includes means for receiving input for setting an input language and an output language, and means for converting the language of the title or description of each of the plurality of steps included in the structured text from the input language to the output language, and provisionally generating the electronic manual based on the structured text and the plurality of sub-videos or still images may include provisionally generating the electronic manual based on the title or description converted into the output language of each of the plurality of steps and the plurality of sub-videos or still images.
[0024] In one aspect of the present invention, the program of the present invention is a program executed in a computer system for assisting in the creation of an electronic manual, the computer system having a processor unit that controls the operation of the computer system, and when the program is executed by the processor unit, the processor unit at least performs the following: receiving one or more videos; receiving information indicating conditions for converting the videos into a plurality of steps; generating structured text for constituting the plurality of steps from audio contained in the one or more videos based on the conditions, wherein the structured text includes at least a title or description of each of the plurality of steps; dividing the one or more videos into a plurality of sub-videos or still images based at least on the one or more videos and the structured text; and provisionally generating the electronic manual based on the structured text and the plurality of sub-videos or still images.
[0025] In one aspect of the present invention, a program of the present invention is a program for assisting in the creation of an electronic manual, the program being executed on a user device having a processor unit that controls the operation of the user device, and when executed by the processor unit, the program causes the processor unit to at least perform the following: identify one or more videos; identify information indicating conditions for converting into a plurality of steps; generate structured text for constituting the plurality of steps from audio contained in the one or more videos based on the conditions, the structured text including at least titles or descriptions of each of the plurality of steps; divide the one or more videos into a plurality of sub-videos or still images based at least on the one or more videos and the structured text; and tentatively generate the electronic manual based on the structured text and the plurality of sub-videos or still images.
[0026] According to the present invention, by providing a computer system and a program for supporting the creation of an electronic manual, it is possible to reduce the time and effort required to create an electronic manual.
[0027] FIG. 1 shows an example of a screen 100 displayed on a user device. FIG. 2 shows an example of a screen 110 displayed on a user device. FIG. 3 shows an example of a screen 120 displayed on a user device. FIG. 4 shows an example of a screen 130 displayed on a user device. FIG. 5 shows an example of the configuration of a system 200 for supporting the creation of an electronic manual. FIG. 6 shows an example of a process executed in a computer system 210. FIG. 7 shows another example of a process executed in the computer system 210.
[0028] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.
[0029] 1. Transition of Screens Displayed on a User Device FIG. 1A shows an example of a screen 100 displayed on a user device. Screen 100 is a screen for specifying one or more videos that will serve as the basis for an electronic manual to be created. Note that screen 100 may be displayed on a user device by having the program of the present invention pre-installed on the user device, or may be displayed on a user device by having the user device communicate with a computer system on which the program of the present invention is pre-installed.
[0030] In the example shown in FIG. 1A , a screen 100 includes a video selection area 101 for selecting one or more videos that will serve as the basis for the electronic manual to be created; an input language setting area 102 for setting the input language of the electronic manual to be created (i.e., the language of the audio included in the one or more videos that will serve as the basis for the electronic manual to be created); an output language setting area 103 for setting the output language of the electronic manual to be created (i.e., the language of the titles and descriptions of each of the multiple steps in the electronic manual to be created); and a transition area 104 for transitioning to the next screen (e.g., screen 110 in FIG. 1B ). When a user selects the video selection area 101, a list of at least one video stored in the memory of the user device is displayed. By selecting one or more of the at least one displayed video, the user can identify one or more videos that will serve as the basis for the electronic manual to be created. In the example shown in FIG. 1A , “Japanese” is selected in the input language setting area 102, and “Japanese” is selected in the output language setting area 103. 1A, a pull-down system is used for the input language setting area 102 and the output language setting area 103, and the user can change the input language of the electronic manual he or she wants to create by selecting the input language setting area 102, and can change the output language of the electronic manual he or she wants to create by selecting the output language setting area 103. Setting the input language in the input language setting area 102 makes it possible to improve the accuracy of the structured text in the subsequent stage of generating the structured text.
[0031] It is possible to transition from screen 100 to the next screen by selecting one or more videos that will form the basis of the electronic manual to be created in video selection area 101, selecting the input language of the electronic manual to be created in input language setting area 102, and selecting the output language of the electronic manual to be created in output language setting area 103, and then selecting transition area 104. Note that transition area 104 may be in a state where it cannot be selected until both the selection of one or more videos that will form the basis of the electronic manual to be created and the selection of the input and output languages of the electronic manual to be created are completed.
[0032] Note that one or more videos selected in the video selection area 101 may include audio. The audio included in one or more videos selected in the video selection area 101 is associated with the time at which the audio is emitted within the playback time of the one or more videos selected in the video selection area 101. If one or more videos selected in the video selection area 101 do not include audio, a warning indicating that one or more videos selected in the video selection area 101 do not include audio may be displayed on the user device after the transition area 104 is selected. At this time, a screen requesting audio input is displayed on the user device, and when the user inputs audio, the screen 100 transitions to the next screen (e.g., screen 110 in FIG. 1B ).
[0033] Furthermore, the language of the audio included in one or more videos selected in the video selection area 101 may be automatically detected. For example, if the input language selected in the input language setting area 102 is different from the automatically detected language of the audio included in one or more videos, a screen requesting the user to confirm the input language may be presented to the user via the user device. This reduces the risk that the input language selected in the input language setting area 102 is different from the language of the audio included in one or more videos, thereby preventing a reduction in the accuracy of the structured text.
[0034] 1B shows an example of a screen 110 displayed on a user device. The screen 110 is a screen for inputting conditions for converting audio included in one or more videos selected in the video selection area 101 into multiple steps. The screen 110 is an example of a screen transitioned to from the screen 100 shown in FIG. 1A when the transition area 104 in the screen 100 shown in FIG. 1A is selected by the user.
[0035] In the example shown in FIG. 1B , the screen 110 includes an area 111 for specifying a "step granularity" related to a limit on the number of steps in the electronic manual, an area 112 for specifying a limit on the number of characters in the title of each step in the electronic manual, an area 113 for specifying a limit on the number of characters in the description of each step in the electronic manual, an area 114 for specifying the wording of the description of each step in the electronic manual, an area 115 for specifying an expected reader of the electronic manual, an area 116 for specifying whether or not subtitles will be included in the electronic manual, and a provisional generation area 117 for performing provisional generation of the electronic manual. In the example shown in FIG. 1B , a pull-down system is used in the area 111, and the "step granularity" can be changed by selecting the area 111. The same applies to each of the areas 112, 113, 114, 115, and 116.
[0036] In the example shown in FIG. 1B, in area 111, "standard" is selected as the "step granularity," in area 112, "up to 30 characters" is selected as the character limit for the title of each step in the electronic manual, in area 113, "approximately 100 characters" is selected as the character limit for the description of each step in the electronic manual, in area 114, "careful" is selected as the wording for the description of each step in the electronic manual, in area 115, "beginners" is selected as the expected reader of the electronic manual, and in area 116, "yes" is selected as the presence or absence of subtitles in the electronic manual.
[0037] In area 111, the "step granularity" may be selected from, for example, dense, standard, sparse, etc., but the present invention is not limited thereto. That is, the "step granularity" may be selected from two or more options. Furthermore, in area 112, the number of characters for the title of each step in the electronic manual may be selected from, for example, up to 10 characters, up to 15 characters, up to 20 characters, up to 30 characters, etc., or from approximately 10 characters, approximately 15 characters, approximately 20 characters, approximately 30 characters, etc., but the present invention is not limited thereto. Furthermore, in area 113, the number of characters for the description of each step in the electronic manual may be selected from, for example, up to 25 characters, up to 50 characters, up to 75 characters, up to 100 characters, up to 125 characters, up to 150 characters, etc., or from approximately 25 characters, approximately 50 characters, approximately 75 characters, approximately 100 characters, approximately 125 characters, approximately 150 characters, etc., but the present invention is not limited thereto. In area 114, the wording of the explanations for each step in the electronic manual can be selected from, for example, polite, frank, etc., but the present invention is not limited to this. In area 115, the expected readers of the electronic manual can be selected from, for example, beginners, intermediate users, advanced users, etc., but the present invention is not limited to this. In area 116, the presence or absence of subtitles in the electronic manual can be selected from, yes or no.
[0038] By selecting the "step granularity" in area 111, selecting the number of characters in the title of each step in the electronic manual in area 112, selecting the number of characters in the explanation of each step in the electronic manual in area 113, selecting the wording of the explanation of each step in the electronic manual in area 114, selecting the expected viewers of the electronic manual in area 115, and selecting whether or not subtitles will be included in the electronic manual in area 116, and then selecting temporary generation area 117, it is possible to provisionally generate an electronic manual based on the "conditions for converting audio included in one or more videos into multiple steps" selected in each of areas 111 to 116, and transition from screen 110 to the next screen. Note that the temporary generation area 117 may be in a state where it cannot be selected until the selections in areas 111 to 116 are completed.
[0039] 1B, an example of selecting the "granularity of steps" in the area 111 has been described, but the present invention is not limited to this. For example, in the area 111, it may be possible to select the number of steps in the electronic manual (e.g., one of 2, 3, 4, 5, 6, 7, 8, 9, and 10).
[0040] 1C shows an example of a screen 120 displayed on a user device. The screen 120 is a preview screen for viewing a provisionally generated electronic manual. The screen 120 is an example of a screen transitioned to from the screen 110 shown in FIG. 1B when the provisionally generated area 117 in the screen 110 shown in FIG. 1B is selected by the user.
[0041] 1C , the screen 120 includes an overview area 121 for providing an overview of the provisionally generated electronic manual, a step area 122 for displaying each of a plurality of steps, a redo area 123 for redoing the provisional generation of the electronic manual, an editing start area 124 for executing editing of the provisionally generated electronic manual, and a final generation area 125 for executing final generation of the provisionally generated electronic manual. The redo area 123, the editing start area 124, and the final generation area 125 are configured to be selectable.
[0042] The summary of the provisionally generated electronic manual displayed in the summary area 121 may be input before the provisional generation of the electronic manual, or may be automatically generated during the provisional generation of the electronic manual. In the example shown in FIG. 1C , the screen 120 displays a first step, a second step, and a portion of a third step among the multiple steps. However, the user can view all of the multiple steps by performing a predetermined operation (e.g., vertical scrolling). When the user selects the redo area 123, the screen 120 transitions to the screen 110 of FIG. 1B, allowing the user to redo the input of the conditions for converting audio included in one or more videos into multiple steps. When the user selects the edit start area 124, the screen 120 transitions to the screen 130 of FIG. 1D, allowing the user to edit the provisionally generated electronic manual. When the user selects the final generation area 125, the final generation of the provisionally generated electronic manual is executed.
[0043] 1C , each step area 122 includes an image area 126 for displaying a sub-movie or a still image divided from one or more movies selected in the movie selection area 101 of FIG. 1A , a title area 127 for displaying the title of the step, and an explanation area 128 for displaying an explanation of the step. The image area 126 may display a movie, as in the image area of the first step, or a still image, as in the image area of the second step. When a movie is displayed in the image area 126, the image area 126 is configured to be selectable, and the movie can be played in response to a user operation (e.g., tapping, clicking, or hovering) for selecting the image area 126.
[0044] The number of steps displayed on screen 120, the number of characters in the title of each step, the number of characters in the explanation for each step, and the wording of the explanation for each step are in accordance with the "conditions for converting audio included in one or more videos into multiple steps" selected in each of areas 111 to 114 on screen 110 in Fig. 1B. Furthermore, the presence or absence of subtitles in the video displayed in image area 126 for each step is in accordance with the "conditions for converting audio included in one or more videos into multiple steps" selected in area 116 on screen 110 in Fig. 1B. The audio included in one or more videos may be in a colloquial style, while the title and explanation for each step may be in a formal style.
[0045] 1D shows an example of a screen 130 displayed on a user device. The screen 130 is a screen for editing a provisionally generated electronic manual. The screen 130 is an example of a screen transitioned to from the screen 120 shown in FIG. 1C when the user selects the editing start area 124 in the screen 120 shown in FIG. 1C.
[0046] 1D , the screen 130 includes a video area 131 for displaying a video, a step sequence area 132 for displaying a sequence of multiple steps, an indicator area 133 for displaying an indicator for editing one or more videos, a division area 134 for dividing one or more videos, and an editing end area 135 for ending editing of the provisionally generated electronic manual. The division area 134 and the editing end area 135 are configured to be selectable.
[0047] The indicator area 133 horizontally represents a timeline of one or more videos selected in the video selection area 101 of Fig. 1A. The left end of the indicator area 133 may be, for example, the playback start time (i.e., 0 minutes 0 seconds) of one or more videos selected in the video selection area 101 of Fig. 1A, and the right end of the indicator area 133 may be, for example, the playback end time (e.g., M minutes S seconds) of one or more videos selected in the video selection area 101 of Fig. 1A. Here, M is an integer from 0 to 59, and S is an integer from 1 to 59.
[0048] In the example shown in FIG. 1D, the indicator area 133 includes a current position indicator 136 indicating the current playback position, a division position indicator 137 indicating the division position of the video that was automatically divided when the provisional generation of the electronic manual was performed, a candidate division time period indicator 138 indicating a candidate division time period between steps of the provisionally generated electronic manual, and a deletion time period indicator 139 indicating a time period of the video that was automatically deleted for a specified reason (e.g., no change in the image appears for a specified time period) when the provisional generation of the electronic manual was performed.
[0049] A video at a playback time corresponding to the location of current position indicator 136 is displayed in video area 131. Current position indicator 136 can slide horizontally on indicator area 133. When a user selects split area 134, a split position indicator 137 can be placed at the position of indicator area 133, and one or more videos can be split at the position of indicator area 133.
[0050] The displayed division position indicator 137 may be, for example, slidable horizontally within the candidate division time slot indicator 138, thereby making it possible to adjust the division position within the candidate division time slot between steps of the provisionally generated electronic manual. Note that the displayed division position indicator 137 may also be slidable horizontally beyond the candidate division time slot indicator 138.
[0051] 1D , the step sequence area 132 includes join indicators 140 for joining adjacent steps. The number of join indicators 140 corresponds to the number of split position indicators 137 displayed in the indicator area 133. The join indicators 140 are configured to be selectable. By selecting a join indicator 140, the selected join indicator 140 disappears, enabling the user to combine two adjacent steps into one step. At this time, the split position indicator 137 corresponding to the disappeared join indicator 140 also disappears.
[0052] The user can restore the automatically deleted video by performing a predetermined operation on the deletion time zone indicator 139 .
[0053] 1A, the multiple videos may be displayed consecutively in the indicator area 133. In this case, a reordering area (not shown) for changing the order of the multiple videos may be displayed on the screen 130, and the order of the multiple videos may be changed in the indicator area 133 according to the user's selection of the reordering area.
[0054] In this way, a user can easily provisionally and finally generate an electronic manual (e.g., a step-structured electronic manual) by selecting one or more videos that will form the basis of the electronic manual and inputting "conditions for converting audio contained in one or more videos into multiple steps." Furthermore, the user can easily edit the provisionally generated electronic manual using the indicator area 133, current position indicator 136, division position indicator 137, candidate division time zone indicator 138, deletion time zone indicator 139, and combination indicator 140 as guides.
[0055] 2. Configuration of a System for Supporting the Creation of an Electronic Manual FIG. 2 shows an example of the configuration of a system 200 for supporting the creation of an electronic manual.
[0056] In the embodiment shown in FIG. 2, the system 200 includes a computer system 210 for supporting the creation of an electronic manual and a user device 220. 1 ~220 N The computer system 210 communicates with a user device 220 via the Internet 230. 1 ~220 N The user device 220 is configured to be able to communicate with each of the 1 ~220 N can be manipulated by a user who wishes to create an electronic manual, where N is an integer greater than or equal to 1.
[0057] Computer system 210 is an information processing system that executes processes for a management company that provides and manages programs for supporting the creation of electronic manuals. In the embodiment shown in Fig. 2, computer system 210 includes an interface unit 211, a processor unit 212 including one or more central processing units (CPUs), and a memory unit 213. The hardware configuration of computer system 210 is not particularly limited as long as it can realize its functions, and may be configured as a single machine or a combination of multiple machines.
[0058] The interface unit 211 is connected to the user device 220 1 ~220 N Controls communication with each of the
[0059] The memory unit 213 stores programs required to execute processing, data required to execute the programs, and the like. Here, it does not matter how the programs are stored in the memory unit 213. For example, the programs may be pre-installed in the memory unit 213. Alternatively, the programs may be installed in the memory unit 213 by being downloaded via a network such as the Internet 230, or may be installed in the memory unit 213 via a storage medium such as an optical disk or a USB.
[0060] The processor unit 212 controls the overall operation of the computer system 210. The processor unit 212 reads and executes programs stored in the memory unit 213. This allows the computer system 210 to function as a device that executes desired steps, and the processor unit 212 of the computer system 210 to operate as a means for achieving desired functions.
[0061] 2, the computer system 210 is connected to a database unit 240. The database unit 240 can store, for example, an electronic manual that has been provisionally generated and then fully generated.
[0062] user device 220 1 is configured to be able to communicate with computer system 210 via Internet 230. In the embodiment shown in FIG. 1 The user device 220 includes an interface unit 221, a processor unit 222, a memory unit 223, a display unit 224, and an input unit 225 for receiving input (e.g., sound, input by selection (e.g., tap, click), etc.). 1The user device 220 may further include, for example, an output unit (not shown) for outputting an output (e.g., sound). 1 The user device 220 may be a mobile wireless terminal such as a mobile phone, a smartphone, or a tablet terminal, or may be a personal computer such as a laptop PC or a notebook PC. 1 The configurations of the interface unit 221, processor unit 222, and memory unit 223 of the user device 220 are similar to those of the interface unit 211, processor unit 212, and memory unit 213 of the computer system 210, and therefore detailed description thereof will be omitted here. The memory unit 223 stores one or more videos that can serve as the basis for an electronic manual. 2 ~220 N The same is true for .
[0063] It should be noted that in the embodiment shown in FIG. 1 ~220 N are described as being capable of communicating with computer system 210 via Internet 230, the invention is not so limited, and any type of network may be used in place of Internet 230.
[0064] 2, the database unit 240 is provided outside the computer system 210, but the present invention is not limited to this. The database unit 240 can also be provided inside the computer system 210. The configuration of the database unit 240 is not limited to a specific hardware configuration. For example, the database unit 240 may be configured as a single hardware component, or may be configured as multiple hardware components. For example, the database unit 240 may be configured as a single external hard disk drive of the computer system 210, or may be configured as cloud storage connected via a network.
[0065] 3. Processing Executed in the Computer System Fig. 3 shows an example of processing executed in the computer system 210. Each step shown in Fig. 3 is executed, for example, by the processor unit 212 of the computer system 210. Each step shown in Fig. 3 will be described below.
[0066] Step S301: One or more videos are identified. The computer system 210 may transmit the one or more videos to, for example, the user device 220. 1 The one or more videos may be received from the user device 220, for example, and may be identified. 1 This is a video that the user who operates the device wants to use as the basis for the electronic manual. This process may correspond to, for example, an operation on the video selection area 101 in FIG. 1A.
[0067] At this time, the computer system 210 may receive input for setting an input language (i.e., the language of the audio included in one or more videos) and an output language (i.e., the language of the electronic manual in the provisional generation and final generation). This process may correspond to, for example, operations on the input language setting area 102 and the output language setting area 103 in FIG. 1A. By receiving an input for setting the input language, the computer system 210 can improve the accuracy of the structured text in step S308. Furthermore, by receiving an input for setting the output language, the computer system 210 can create an electronic manual in a language that is the same as the input language or a language different from the input language.
[0068] Step S302: It is determined whether the one or more videos received in step S301 contain audio. The audio contained in the one or more videos may be audio indicating procedures in an electronic manual. If the determination result is "Yes," the process proceeds to step S307. If the determination result is "No," the process proceeds to step S303.
[0069] Step S303: A process is performed to warn the user that one or more videos received in step S301 do not contain audio. This process may be performed, for example, by the computer system 210 sending a warning to the user device 220 indicating that one or more videos do not contain audio. 1 and sends the alert to the user device 220 1 This may be accomplished by the above presentation, or by the computer system 210 transmitting an audible warning signal to the user device 220 indicating that one or more videos do not contain audio. 1 and transmits the warning sound to the user device 220 1 This may be achieved by issuing the above.
[0070] Step S304: It is determined whether a user input indicating that a voice is to be input is received. The user input indicating that a voice is to be input is, for example, received from the user device 220. 1 If the determination result is "Yes", the process proceeds to step S306, and if the determination result is "No", the process proceeds to step S305.
[0071] Step S305: A process is executed to notify the user that an electronic manual cannot be created. This process is carried out, for example, by the computer system 210 transmitting information indicating that an electronic manual cannot be created to the user device 220. 1 and transmits the information to the user device 220 1 This may be achieved by the method presented above.
[0072] Step S306: It is determined whether or not a voice input is received. The voice input may be, for example, received from the user device 220. 1 The audio input may be achieved, for example, by inputting pre-recorded audio or by transmitting one or more videos to the user device 220. 1 This may be achieved by recording the audio in parallel with playing it on the screen. If the determination result is "Yes", the process proceeds to step S307, and if the determination result is "No", the process returns to step S306.
[0073] Step S307: Conditions for converting audio included in one or more videos into a plurality of steps are identified. The conditions for converting into a plurality of steps include at least a limit on the number of steps, which may correspond to the operation for area 111 in FIG. 1B . The conditions for converting into a plurality of steps may further include a limit on the number of characters in the title (e.g., a limit on the number of characters in the title of each step in the electronic manual) and / or a limit on the number of characters in the description (e.g., a limit on the number of characters in the description of each step in the electronic manual), which may correspond to the operation for areas 112 and 113 in FIG. 1B . The conditions for converting into a plurality of steps may further include a limit on the wording of the description (e.g., a limit on the wording of the description of each step in the electronic manual) and / or an expected reader of the electronic manual, which may correspond to the operation for areas 114 and 115 in FIG. 1B .
[0074] Step S308: Structured text for constituting the multiple steps of the electronic manual is generated. The structured text is generated from the audio included in the one or more videos identified in step S301 based on conditions for converting the audio included in the one or more videos into multiple steps. The structured text includes at least a title or description for each of the multiple steps. The title of each of the multiple steps included in the structured text may correspond, for example, to the description in the title area 127 of the screen 120 in FIG. 1C. The description of each of the multiple steps included in the structured text may correspond, for example, to the description in the description area 128 of the screen 120 in FIG. 1C. The structured text may be generated using, for example, artificial intelligence (e.g., ChatGPT). The computer system 210 is configured to convert the structured text (particularly, the title or description for each of the multiple steps included in the structured text) from an input language to an output language. This allows the computer system 210 to generate the structured text in the set output language even if the input language of the audio included in the one or more videos is different from the output language of the electronic manual.
[0075] Note that computer system 210 may generate the structured text directly from the audio included in one or more videos based on the conditions for converting the audio included in one or more videos into a plurality of steps. Alternatively, computer system 210 may convert the audio included in one or more videos into text by transcribing the audio included in one or more videos, and generate the structured text based on the converted text and the conditions for converting the audio included in one or more videos into a plurality of steps.
[0076] Step S309: One or more videos are divided into multiple sub-videos or still images. This process is performed based at least on the one or more videos identified in step S301 and the structured text generated in step S308. This process can be achieved by the computer system 210, for example, identifying scene change timings within the videos based on the structured text and generating multiple sub-videos or still images by dividing the one or more videos based on the scene change timings. The identification of the scene change timings may be achieved, for example, by identifying breaks in the content of the structured text based on the structured text and identifying timings in the audio corresponding to the breaks in the structured text as the scene change timings. The breaks in the content of the structured text may exist, for example, between steps of multiple steps.
[0077] The computer system 210 may generate multiple sub-movies or still images from one or more videos by, for example, identifying timings of significant image changes in one or more videos, identifying timings of audio breaks, and dividing the one or more videos at timings where the timings of significant image changes, scene changes, and audio breaks coincide. The timings of significant image changes in one or more videos may be, for example, timings where the area of the image that has changed relative to the display area of the video exceeds a predetermined threshold. The timings of audio breaks may be, for example, timings where a period of silence in one or more videos exists for a predetermined length of time.
[0078] The computer system 210 may accomplish segmenting the one or more videos into multiple sub-videos or still images by, for example, segmenting the one or more videos into multiple candidate sub-videos based at least on the one or more videos and the structured text, identifying a candidate sub-video among the multiple candidate sub-videos that exhibits audio above a predetermined volume for a predetermined time while exhibiting no image change, and converting the candidate sub-video into a still image based on the candidate sub-video. Converting the candidate sub-video into a still image may be accomplished, for example, by capturing a portion of the candidate sub-video as a still image.
[0079] Step S310: A provisional electronic manual is generated. This process is executed based on the structured text generated in step S308 and the plurality of subsidiary moving images or still images generated in step S309. This process is executed by receiving a user input requesting provisional generation of an electronic manual from the user device 220. 11B , for example. This process may correspond to an operation on the provisional generation area 117 of FIG. 1B . The provisionally generated electronic manual may not include the audio contained in one or more videos. The language of the provisionally generated electronic manual may be changed from the input language to the output language depending on the language settings in the input language setting area 102 and the output language setting area 103 of FIG. 1A . If the input language of the audio contained in one or more videos is different from the output language of the electronic manual, the language of the provisionally generated electronic manual may be changed, for example, by machine translation. Furthermore, if the structured text (especially the titles or descriptions of each of the multiple steps included in the structured text) has been converted into the output language, the computer system 210 can provisionally generate the electronic manual based on the titles or descriptions of each of the multiple steps converted into the output language and the multiple sub-videos or still images generated in step S309.
[0080] Step S311: It is determined whether a user input for executing the book generation of the electronic manual is received. The user input for executing the book generation of the electronic manual is received from, for example, the user device 220. 1 If the determination result is "Yes", the process proceeds to step S312, and if the determination result is "No", the process returns to step S311.
[0081] Step S312: Final generation of the electronic manual is executed. This completes the electronic manual. The final generated electronic manual may be output according to the language setting in the output language setting area 103 of FIG. 1A. When final generation of the electronic manual is executed, the computer system 210 may generate audio data for reading the structured text generated in step S308. This makes it possible to realize automatic reading of the completed electronic manual. Furthermore, by generating audio data in multiple languages for reading the structured text generated in step S308, it is possible to provide the electronic manual in multiple languages regardless of the language of the audio included in one or more videos.
[0082] Figure 4 shows another example of processing executed in the computer system 210. Each step shown in Figure 4 is executed, for example, by the processor unit 212 of the computer system 210. Each step shown in Figure 4 shows an example of processing for editing a provisionally generated electronic manual at any timing after step S311 and before step S312 in Figure 3. Each step shown in Figure 4 will be described below.
[0083] Step S401: A user input indicating a desire to edit the provisionally generated electronic manual is received. The user input indicating a desire to edit the provisionally generated electronic manual is received, for example, from the user device 220. 1 This process may correspond to, for example, an operation on the edit start area 124 in FIG. 1C.
[0084] Step S402: A division candidate time period between the steps of the provisionally generated electronic manual is identified. Within the division candidate time period, the user may be able to adjust the division position between the steps of the provisionally generated electronic manual. The division candidate time period is identified, for example, based on the structured text generated in step S308 and the audio included in one or more videos. Specifically, the computer system 210 may, for example, identify the audio playback time corresponding to each of the steps based on the structured text generated in step S308, and identify the division candidate time period based on the audio playback time corresponding to each step. For example, the division candidate time period may be the entire time period between the end of the audio playback time corresponding to a step and the start of the audio playback time corresponding to the next step, or may be a time period within a predetermined range from a certain point between the end of the audio playback time corresponding to a step and the start of the audio playback time corresponding to the next step.
[0085] Step S403: A process is executed to present candidate time slots for division between the steps of the provisionally generated electronic manual. This process is performed, for example, by the computer system 210 transmitting information indicating candidate time slots for division between the steps of the provisionally generated electronic manual to the user device 220. 1 and transmits the information to the user device 220 1 This process may be accomplished by, for example, displaying the image 130 of FIG. 1 This allows the user to start editing the provisionally generated electronic manual.
[0086] Step S404: It is determined whether a user input for editing the provisionally generated electronic manual is received. The user input for editing the provisionally generated electronic manual may be received from, for example, the user device 220. 1 This process may correspond to, for example, an operation on the image 130 in Fig. 1D. If the determination result is "Yes", the process proceeds to step S405, and if the determination result is "No", the process proceeds to step S406.
[0087] The user input for editing the provisionally generated electronic manual may include, for example, a user input for adjusting a division position between steps of the provisionally generated electronic manual within a division candidate time period. This may correspond to an operation on the division position indicator 137 in FIG. 1D . The user input for editing the provisionally generated electronic manual may also include, for example, a user input for dividing one or more videos. This may correspond to an operation on the division area 134 in FIG. 1D . The user input for editing the provisionally generated electronic manual may also include a user input for restoring automatically deleted videos. This may correspond to an operation on the deletion time period indicator 139 in FIG. 1D . The user input for editing the provisionally generated electronic manual may also include a user input for merging adjacent steps. This may correspond to an operation on the merge indicator 140 in FIG. 1D .
[0088] Step S405: Editing of the temporarily generated electronic manual is performed in accordance with a user input for editing the temporarily generated electronic manual. For example, if the user input for editing the temporarily generated electronic manual is a user input for adjusting the division positions between the multiple steps of the temporarily generated electronic manual within the time period of the division candidate, adjustment of the division positions between the multiple steps of the temporarily generated electronic manual within the time period of the division candidate is performed.
[0089] Step S406: It is determined whether a user input for terminating the editing of the provisionally generated electronic manual has been received. The user input for terminating the editing of the provisionally generated electronic manual may be received from, for example, the user device 220. 1 This process may correspond to, for example, an operation on the edit end area 135 in FIG.
[0090] An example of generating structured text using ChatGPT will be described below.
[0091] For example, suppose the audio included in one or more videos is "First, open the Settings app. After opening the Settings app, open General management at the bottom of the left menu. Then, tap Text-to-speech in the right menu. Next, tap the gear icon next to Priority Engine. The already installed languages are listed at the bottom right. If your language is not there, tap Install Audio Data. If you tap a language that is not installed, a screen like this will appear prompting you to install it, so press the Install button. When the installation is complete, a screen like this will appear notifying you of completion. This completes the audio data installation method." Using ChatGPT under the conditions of "up to 50 steps," "up to 50 characters for the title," and "up to 200 characters for the description," the following output sentence can be output as structured text from this audio:
[0092] In the above output sentence, "Step 0: ..." represents the title of each step, and "Description" represents the description of each step. Also, "Base Text" refers to the text that is generated from the audio corresponding to each step. In the above audio example, the structured text includes seven steps.
[0093] When generating the structured text, a correspondence between each step of the multiple steps and the playback time of the audio corresponding to each step is maintained and / or recorded. In the audio example described above, the base text of step 1, "First, open the Settings app.", corresponds to the playback time of the audio included in one or more videos from 0 minutes 0 seconds to 0 minutes 18 seconds, the base text of step 2, "Open General management at the bottom of the left menu.", corresponds to the playback time of the audio from 0 minutes 19 seconds to 0 minutes 23 seconds, the base text of step 3, "Tap Read Text to Speech.", corresponds to the playback time of the audio from 0 minutes 24 seconds to 0 minutes 30 seconds, the base text of step 4, "Tap the gear icon next to the priority engine.", corresponds to the playback time of the audio from 0 minutes 32 seconds to 0 minutes 38 seconds, and the base text of step 5, "Already installed.", corresponds to the playback time of the audio included in one or more videos from 0 minutes 0 seconds to 0 minutes 38 seconds. The installed languages are listed at the bottom right, and if your language is not listed, tap Install audio data. Then, tap a language that is not installed" corresponds to the audio playback time from 0 minutes 40 seconds to 0 minutes 48 seconds, the base text of step 6 "When you tap a language that is not installed, a screen like this will appear prompting you to install it, so press the Install button" corresponds to the audio playback time from 0 minutes 50 seconds to 1 minute 02 seconds, and the base text of step 7 "When the installation is complete, a screen like this will appear notifying you of the completion" corresponds to the audio playback time from 1 minute 04 seconds to 1 minute 15 seconds. In this case, for example, since step 1 of the structured text corresponds to the audio playback time of "0 minutes 0 seconds to 0 minutes 18 seconds," and step 2 of the structured text corresponds to the audio playback time of "0 minutes 19 seconds to 0 minutes 23 seconds," computer system 210 can identify the candidate division time period between step 1 and step 2 as "0 minutes 18 seconds to 0 minutes 19 seconds," which is between "0 minutes 0 seconds to 0 minutes 18 seconds" and "0 minutes 19 seconds to 0 minutes 23 seconds." When provisional generation of an electronic manual is performed, the division position of the video between step 1 and step 2 can be determined within the candidate division time period of "0 minutes 18 seconds to 0 minutes 19 seconds."The same applies to the time period of the splitting candidates between steps 2 and 3, the time period of the splitting candidates between steps 3 and 4, the time period of the splitting candidates between steps 4 and 5, the time period of the splitting candidates between steps 5 and 6, and the time period of the splitting candidates between steps 6 and 7.
[0094] When the computer system 210 determines the division position of the video between step 1 and step 2 within the division candidate time period "0:18 to 0:19", it may, for example, automatically determine the division position to be in the center of the division candidate time period "0:18 to 0:19", or may determine the division position randomly within the division candidate time period "0:18 to 0:19". The same applies to the other division candidate time periods.
[0095] 3 and 4, an example in which the computer system 210 executes the processes of the steps shown in FIGS. 3 and 4 has been described, but the present invention is not limited to this. For example, the processes of the steps shown in FIGS. 3 and 4 may be executed by, for example, the user device 220 instead of the computer system 210. 1 (In particular, user device 220 1 In this case, the user device 220 1 In step S301 of FIG. 3, the user device 220 1 By selecting one or more of the multiple videos stored in the memory unit 223 in the video selection area 101 of Fig. 1A, it is possible to specify one or more videos that will be the basis of the electronic manual, and by selecting "conditions for converting audio included in one or more videos into multiple steps" in each of the areas 111 to 116 of Fig. 1B in step S307 of Fig. 3, it is possible to specify the conditions for converting audio included in one or more videos into multiple steps. 1 In steps S304 and S311 of FIG. 3 and step S406 of FIG. 4, the user device 220 1 The input 225 receives user input.
[0096] 3 and 4, an example has been described in which the processing of each step shown in Fig. 3 and 4 is realized by the processor executing a program stored in the memory, but the present invention is not limited to this. At least some of the processing of each step shown in Fig. 3 and 4 may be realized by a hardware configuration such as a control circuit.
[0097] As described above, the present invention has been illustrated using the preferred embodiment of the present invention, but the present invention should not be interpreted as being limited to this embodiment. It is understood that the scope of the present invention should be interpreted only by the claims. It is understood that a person skilled in the art can implement an equivalent scope based on the description of the present invention and common technical knowledge from the description of the specific preferred embodiment of the present invention.
[0098] The present invention is useful for reducing the time and effort required to create an electronic manual by providing a computer system, a program, etc. for supporting the creation of an electronic manual.
[0099] 200 System 210 Computer System 220 1 ~220 N User device 230 Internet 240 Database unit
Claims
1. A computer system for assisting in the creation of an electronic manual, comprising: means for receiving one or more videos; means for receiving information indicating conditions for converting the videos into a plurality of steps; means for generating structured text for constituting the plurality of steps from audio included in the one or more videos based on the conditions, the structured text including at least a title or description of each of the plurality of steps; means for dividing the one or more videos into a plurality of sub-videos or still images based at least on the one or more videos and the structured text; and means for provisionally generating the electronic manual based on the structured text and the plurality of sub-videos or still images.
2. The computer system of claim 1, wherein the conditions include a limit on the number of steps.
3. The computer system of claim 2, wherein the conditions further include a limit on the number of characters in the title and / or a limit on the number of characters in the description.
4. The computer system according to claim 1, wherein the audio included in the one or more videos is audio that indicates the procedures in the electronic manual.
5. The computer system according to claim 4, wherein the provisionally generated electronic manual does not include audio included in the one or more videos.
6. The computer system of claim 1, wherein dividing the one or more videos into a plurality of sub-videos or still images based at least on the one or more videos and the structured text comprises: dividing the one or more videos into a plurality of candidate sub-videos based at least on the one or more videos and the structured text; and converting, from among the plurality of candidate sub-videos, a candidate sub-video that has audio exceeding a predetermined volume for a predetermined period of time but does not show any image change into a still image based on the candidate sub-video.
7. The computer system of claim 1, wherein dividing the one or more videos into a plurality of sub-videos or still images based at least on the one or more videos and the structured text comprises: identifying timings of scene changes based on the structured text; and generating the plurality of sub-videos or still images by dividing the one or more videos based on the timings of the scene changes.
8. The computer system of claim 7, wherein identifying the timing of the scene change based on the structured text includes: identifying a break in the content of the structured text based on the structured text; and identifying a timing in the audio corresponding to the break in the structured text as the timing of the scene change.
9. The computer system of claim 7, wherein dividing the one or more videos into a plurality of sub-videos or still images based at least on the one or more videos and the structured text further comprises: identifying timings of large image changes in the one or more videos; identifying timings of audio breaks; and generating the plurality of sub-videos or still images by dividing the one or more videos at timings where the timings of large image changes, the timings of the scene changes, and the timings of the audio breaks coincide.
10. The computer system of claim 1, wherein generating the structured text from the audio contained in the one or more videos based on the conditions comprises: converting the audio into text by transcribing the audio contained in the one or more videos; and generating the structured text based on the text converted from the audio and the conditions.
11. The computer system of claim 1, further comprising: means for receiving a first user input indicating a desire to edit the provisionally generated electronic manual; means for identifying candidate division time periods between steps of the provisionally generated electronic manual in response to receiving the first user input, wherein a user can adjust the division positions between steps of the provisionally generated electronic manual within the candidate division time periods; means for presenting the candidate division time periods; means for receiving a second user input for adjusting the division positions between steps of the provisionally generated electronic manual within the candidate division time periods; and means for editing the provisionally generated electronic manual based on the second user input.
12. The computer system of claim 11, wherein identifying the candidate segmentation time periods includes identifying the candidate segmentation time periods based on the structured text and audio included in the one or more videos.
13. The computer system of claim 11, wherein identifying the time periods of the segmentation candidates based on the structured text and the audio included in the one or more videos includes: identifying a playback time of the audio corresponding to each step of the plurality of steps based on the structured text; and identifying the time periods of the segmentation candidates based on the playback time of the audio corresponding to each step.
14. The computer system of claim 1, further comprising: means for receiving a third user input for executing book generation of the electronic manual; and means for executing book generation of the electronic manual in response to receiving the third user input.
15. The computer system of claim 1, further comprising: means for determining whether the one or more videos contain audio; and means for warning a user that the one or more videos do not contain audio if it is determined that the one or more videos do not contain audio.
16. The computer system of claim 1, wherein the audio included in the one or more videos is in a colloquial style, and the title and description are in a written style.
17. The computer system of claim 1, further comprising means for generating audio data for reading the structured text aloud.
18. The computer system of claim 1, further comprising: means for receiving an input for setting an input language and an output language; and means for converting the language of the title or the description of each of the plurality of steps included in the structured text from the input language to the output language; and provisionally generating the electronic manual based on the structured text and the plurality of sub-videos or still images includes provisionally generating the electronic manual based on the title or the description converted into the output language of each of the plurality of steps and the plurality of sub-videos or still images.
19. A program executed in a computer system for supporting the creation of an electronic manual, the computer system having a processor unit that controls the operation of the computer system, the program, when executed by the processor unit, causes the processor unit to at least perform the following: receive one or more videos; receive information indicating conditions for converting into a plurality of steps; generate structured text for constituting the plurality of steps from audio included in the one or more videos based on the conditions, the structured text including at least a title or description of each of the plurality of steps; divide the one or more videos into a plurality of sub-videos or still images based at least on the one or more videos and the structured text; and tentatively generate the electronic manual based on the structured text and the plurality of sub-videos or still images.
20. A program for assisting in the creation of an electronic manual, the program being executed on a user device having a processor unit that controls the operation of the user device, the program causing the processor unit, when executed by the processor unit, to at least perform the following: identify one or more videos; identify information indicating conditions for converting into a plurality of steps; generate structured text for constituting the plurality of steps from audio contained in the one or more videos based on the conditions, the structured text including at least a title or description of each of the plurality of steps; divide the one or more videos into a plurality of sub-videos or still images based at least on the one or more videos and the structured text; and tentatively generate the electronic manual based on the structured text and the plurality of sub-videos or still images.
Citation Information
Patent Citations
Image layout apparatus. image layout program, and image layout method
JP2004120127A
Program, method and device for supporting creation of document
JP2019020789A
Explicit knowledge formalization system and method of the same
JP2019144822A
Video manual creation device, video manual creation method, and video manual creation program
JP7023427B1