Computer system and program for supporting creation of electronic manual

The computer system and program streamline the creation of electronic manuals by converting video content into structured text and images, addressing the time-consuming nature of manual creation.

JP2025127186AActive Publication Date: 2025-09-01STUDIST CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024023759
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-20
Publication Date
2025-09-01
Estimated Expiration
2044-02-20

AI Technical Summary

Technical Problem

Creating an electronic manual, especially one that includes video content, requires significant time and effort.

Method used

A computer system and program that receive videos, convert them into structured text and images, and generate an electronic manual based on conditions such as step limits and language settings, allowing for provisional and final manual creation with user editing capabilities.

Benefits of technology

Reduces the time and effort required to create an electronic manual by automating the conversion and editing process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025127186000001_ABST
    Figure 2025127186000001_ABST
Patent Text Reader

Abstract

To provide a computer system for supporting creation of an electronic manual.SOLUTION: A computer system includes: means for receiving one or more videos; means for receiving information representing a condition for converting audio into a plurality of steps; means for generating structured text for composing the plurality of steps from the audio included in the one or more videos on the basis of the condition, the structured text including at least a title or a description of each of the plurality of steps; means for dividing the one or more videos into a plurality of sub-videos or still images on the basis of at least the one or more videos and the structured text; and means for tentatively generating an electronic manual on the basis of the structured text and the plurality of sub-videos or still images.SELECTED DRAWING: Figure 1C
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a computer system and a program for supporting the creation of an electronic manual. [Background technology]

[0002] BACKGROUND ART It has been known for some time that electronic manuals are created and used for the purpose of improving work efficiency and the like (see, for example, Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] International Publication No. 2017 / 183064 Summary of the Invention [Problem to be solved by the invention]

[0004] However, creating an electronic manual still requires time and effort, and creating an electronic manual that includes video in particular requires a considerable amount of time and effort.

[0005] The present invention has been made in consideration of the above-mentioned problems, and aims to reduce the time and effort required to create an electronic manual by providing a computer system and program for assisting in the creation of an electronic manual. [Means for solving the problem]

[0006] In one aspect of the present invention, a computer system of the present invention is a computer system for assisting in the creation of an electronic manual, and the computer system includes: means for receiving one or more videos; means for receiving information indicating conditions for converting the videos into a plurality of steps; means for generating structured text for constituting the plurality of steps from audio contained in the one or more videos based on the conditions, the structured text including at least a title or description of each of the plurality of steps; means for dividing the one or more videos into a plurality of sub-videos or still images based at least on the one or more videos and the structured text; and means for provisionally generating the electronic manual based on the structured text and the plurality of sub-videos or still images.

[0007] In one embodiment of the present invention, the conditions may include a limit on the number of steps.

[0008] In one embodiment of the present invention, the conditions may further include a limit on the number of characters in the title and / or a limit on the number of characters in the description.

[0009] In one embodiment of the present invention, the audio included in the one or more videos may be audio that indicates the steps of the electronic manual.

[0010] In one embodiment of the present invention, the provisionally generated electronic manual may not include audio included in the one or more videos.

[0011] In one embodiment of the present invention, dividing the one or more videos into a plurality of sub-videos or still images based at least on the one or more videos and the structured text may include dividing the one or more videos into a plurality of candidate sub-videos based at least on the one or more videos and the structured text, and converting a candidate sub-video among the plurality of candidate sub-videos that has audio exceeding a predetermined volume for a predetermined period of time but does not show any change in image into a still image based on the candidate sub-video.

[0012] In one embodiment of the present invention, dividing the one or more videos into a plurality of sub-videos or still images based at least on the one or more videos and the structured text may include identifying timing of scene changes based on the structured text, and generating the plurality of sub-videos or still images by dividing the one or more videos based on the timing of the scene changes.

[0013] In one embodiment of the present invention, identifying the timing of the scene change based on the structured text may include identifying a break in the content of the structured text based on the structured text, and identifying a timing in the audio corresponding to the break in the structured text as the timing of the scene change.

[0014] In one embodiment of the present invention, dividing the one or more videos into multiple sub-videos or still images based at least on the one or more videos and the structured text may further include identifying timings of large image changes in the one or more videos, identifying timings of audio breaks, and generating the multiple sub-videos or still images by dividing the one or more videos at timings where the timings of large image changes, the timings of the scene changes, and the timings of the audio breaks coincide.

[0015] In one embodiment of the present invention, generating the structured text from audio contained in the one or more videos based on the conditions may include converting the audio into text by transcribing the audio contained in the one or more videos, and generating the structured text based on the text converted from the audio and the conditions.

[0016] In one embodiment of the present invention, the computer system may further include means for receiving a first user input indicating a desire to edit the provisionally generated electronic manual; means for identifying, in response to receiving the first user input, a candidate division time period between steps of the provisionally generated electronic manual, wherein within the candidate division time period, a user can adjust the division position between steps of the provisionally generated electronic manual; means for presenting the candidate division time period; means for receiving a second user input for adjusting the division position between steps of the provisionally generated electronic manual within the candidate division time period; and means for editing the provisionally generated electronic manual based on the second user input.

[0017] In one embodiment of the present invention, identifying the time periods of the segmentation candidates may include identifying the time periods of the segmentation candidates based on the structured text and audio contained in the one or more videos.

[0018] In one embodiment of the present invention, identifying the time periods of the division candidates based on the structured text and the audio contained in the one or more videos may include identifying, based on the structured text, a playback time of the audio corresponding to each step of the plurality of steps, and identifying the time periods of the division candidates based on the playback time of the audio corresponding to each step.

[0019] In one embodiment of the present invention, the computer system may further include means for receiving a third user input for executing book generation of the electronic manual, and means for executing book generation of the electronic manual in response to receiving the third user input.

[0020] In one embodiment of the present invention, the computer system may further include means for determining whether the one or more videos contain audio, and means for warning a user that the one or more videos do not contain audio if it is determined that the one or more videos do not contain audio.

[0021] In one embodiment of the present invention, the audio included in the one or more videos may be in a colloquial style, and the title and description may be in a written style.

[0022] In one embodiment of the present invention, the computer system may further comprise means for generating audio data for reading the structured text aloud.

[0023] In one embodiment of the present invention, the computer system includes means for receiving input for setting an input language and an output language, and means for converting the language of the title or description of each of the plurality of steps included in the structured text from the input language to the output language, and provisionally generating the electronic manual based on the structured text and the plurality of sub-videos or still images may include provisionally generating the electronic manual based on the title or description converted into the output language of each of the plurality of steps and the plurality of sub-videos or still images.

[0024] In one aspect of the present invention, the program of the present invention is a program executed in a computer system for assisting in the creation of an electronic manual, the computer system having a processor unit that controls the operation of the computer system, and when the program is executed by the processor unit, the processor unit at least performs the following: receiving one or more videos; receiving information indicating conditions for converting the videos into a plurality of steps; generating structured text for constituting the plurality of steps from audio contained in the one or more videos based on the conditions, wherein the structured text includes at least a title or description of each of the plurality of steps; dividing the one or more videos into a plurality of sub-videos or still images based at least on the one or more videos and the structured text; and provisionally generating the electronic manual based on the structured text and the plurality of sub-videos or still images.

[0025] In one aspect of the present invention, a program of the present invention is a program for assisting in the creation of an electronic manual, the program being executed on a user device, the user device having a processor unit that controls the operation of the user device, and when executed by the processor unit, the program causes the processor unit to at least perform the following: identify one or more videos; identify information indicating conditions for converting into a plurality of steps; generate structured text for constituting the plurality of steps from audio contained in the one or more videos based on the conditions, the structured text including at least titles or descriptions of each of the plurality of steps; divide the one or more videos into a plurality of sub-videos or still images based at least on the one or more videos and the structured text; and tentatively generate the electronic manual based on the structured text and the plurality of sub-videos or still images. [Effects of the Invention]

[0026] According to the present invention, by providing a computer system and a program for supporting the creation of an electronic manual, it is possible to reduce the time and effort required to create an electronic manual. [Brief explanation of the drawings]

[0027] [Figure 1A] FIG. 1 is a diagram showing an example of a screen 100 displayed on a user device. [Figure 1B] FIG. 10 is a diagram showing an example of a screen 110 displayed on a user device. [Figure 1C] FIG. 10 is a diagram showing an example of a screen 120 displayed on a user device. [Figure 1D] FIG. 10 is a diagram showing an example of a screen 130 displayed on a user device. [Figure 2] FIG. 1 shows an example of the configuration of a system 200 for supporting the creation of electronic manuals. [Figure 3] FIG. 2 is a diagram illustrating an example of processing executed in a computer system 210. [Figure 4] FIG. 2 is a diagram illustrating another example of processing executed in the computer system 210. DETAILED DESCRIPTION OF THE INVENTION

[0028] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.

[0029] 1. Screen transitions displayed on the user device 1A shows an example of a screen 100 displayed on a user device. The screen 100 is a screen for specifying one or more videos that will serve as the basis for an electronic manual that the user wants to create. Note that the screen 100 may be displayed on the user device by having the program of the present invention pre-installed on the user device, or may be displayed on the user device by having the user device communicate with a computer system on which the program of the present invention is pre-installed.

[0030] In the example shown in FIG. 1A, a screen 100 includes a video selection area 101 for selecting one or more videos that will serve as the basis for the electronic manual to be created; an input language setting area 102 for setting the input language of the electronic manual to be created (i.e., the language of the audio included in the one or more videos that will serve as the basis for the electronic manual to be created); an output language setting area 103 for setting the output language of the electronic manual to be created (i.e., the language of the titles and descriptions of each of the multiple steps in the electronic manual to be created); and a transition area 104 for transitioning to the next screen (e.g., screen 110 in FIG. 1B). When a user selects the video selection area 101, a list of at least one video stored in the memory of the user device is displayed. By selecting one or more of the at least one displayed video, the user can identify one or more videos that will serve as the basis for the electronic manual to be created. In the example shown in FIG. 1A, “Japanese” is selected in the input language setting area 102, and “Japanese” is selected in the output language setting area 103. 1A, a pull-down system is used for input language setting area 102 and output language setting area 103, and the user can change the input language of the electronic manual he or she wants to create by selecting input language setting area 102, and can change the output language of the electronic manual he or she wants to create by selecting output language setting area 103. Setting the input language in input language setting area 102 makes it possible to improve the accuracy of structured text in the subsequent stage of generating structured text.

[0031] It is possible to transition from screen 100 to the next screen by selecting one or more videos that will be the basis for the electronic manual you want to create in video selection area 101, selecting the input language for the electronic manual you want to create in input language setting area 102, and selecting the output language for the electronic manual you want to create in output language setting area 103, and then selecting transition area 104. Note that transition area 104 may be in a state where it cannot be selected until both the selection of one or more videos that will be the basis for the electronic manual you want to create and the selection of the input and output languages ​​for the electronic manual you want to create are complete.

[0032] Note that one or more videos selected in the video selection area 101 may include audio. The audio included in one or more videos selected in the video selection area 101 is associated with the time at which the audio is emitted within the playback time of the one or more videos selected in the video selection area 101. If one or more videos selected in the video selection area 101 do not include audio, after the transition area 104 is selected, a warning may be displayed on the user device to indicate that one or more videos selected in the video selection area 101 do not include audio. At this time, a screen requesting audio input is displayed on the user device, and when the user inputs audio, the screen 100 transitions to the next screen (e.g., screen 110 in FIG. 1B).

[0033] Furthermore, the language of the audio included in one or more videos selected in the video selection area 101 may be automatically detected. For example, if the input language selected in the input language setting area 102 is different from the automatically detected language of the audio included in one or more videos, a screen requesting the user to confirm the input language may be presented to the user via the user device. This reduces the risk that the input language selected in the input language setting area 102 is different from the language of the audio included in one or more videos, thereby preventing a reduction in the accuracy of the structured text.

[0034] 1B shows an example of a screen 110 displayed on a user device. Screen 110 is a screen for inputting conditions for converting audio contained in one or more videos selected in video selection area 101 into multiple steps. Screen 110 is an example of a screen transitioned to from screen 100 shown in FIG. 1A when transition area 104 in screen 100 shown in FIG. 1A is selected by the user.

[0035] In the example shown in FIG. 1B , screen 110 includes an area 111 for specifying a "step granularity" related to a limit on the number of steps in the electronic manual, an area 112 for specifying a limit on the number of characters in the title of each step in the electronic manual, an area 113 for specifying a limit on the number of characters in the description of each step in the electronic manual, an area 114 for specifying the wording of the description of each step in the electronic manual, an area 115 for specifying the intended audience of the electronic manual, an area 116 for specifying whether or not subtitles will be included in the electronic manual, and a provisional generation area 117 for performing provisional generation of the electronic manual. In the example shown in FIG. 1B , area 111 employs a pull-down system, and the "step granularity" can be changed by selecting area 111. The same applies to each of areas 112, 113, 114, 115, and 116.

[0036] In the example shown in FIG. 1B, in area 111, "standard" is selected as the "step granularity," in area 112, "up to 30 characters" is selected as the character limit for the title of each step in the electronic manual, in area 113, "approximately 100 characters" is selected as the character limit for the description of each step in the electronic manual, in area 114, "careful" is selected as the wording for the description of each step in the electronic manual, in area 115, "beginners" is selected as the expected reader of the electronic manual, and in area 116, "yes" is selected as the presence or absence of subtitles in the electronic manual.

[0037] In area 111, the "step granularity" may be selected from, for example, dense, standard, sparse, etc., but the present invention is not limited thereto. That is, the "step granularity" may be selected from two or more options. In area 112, the number of characters for the title of each step in the electronic manual may be selected from, for example, up to 10 characters, up to 15 characters, up to 20 characters, up to 30 characters, etc., or from approximately 10 characters, approximately 15 characters, approximately 20 characters, approximately 30 characters, etc., but the present invention is not limited thereto. In area 113, the number of characters for the description of each step in the electronic manual may be selected from, for example, up to 25 characters, up to 50 characters, up to 75 characters, up to 100 characters, up to 125 characters, up to 150 characters, etc., or from approximately 25 characters, approximately 50 characters, approximately 75 characters, approximately 100 characters, approximately 125 characters, approximately 150 characters, etc., but the present invention is not limited thereto. In area 114, the wording of the explanations for each step in the electronic manual can be selected from, for example, polite, frank, etc., but the present invention is not limited to this. In area 115, the expected readers of the electronic manual can be selected from, for example, beginners, intermediate users, advanced users, etc., but the present invention is not limited to this. In area 116, the presence or absence of subtitles in the electronic manual can be selected from, yes or no.

[0038] By selecting the "step granularity" in area 111, selecting the number of characters in the title of each step in the electronic manual in area 112, selecting the number of characters in the explanation of each step in the electronic manual in area 113, selecting the wording of the explanation of each step in the electronic manual in area 114, selecting the expected viewers of the electronic manual in area 115, and selecting whether or not subtitles will be included in the electronic manual in area 116, and then selecting temporary generation area 117, it is possible to temporarily generate an electronic manual based on the "conditions for converting audio included in one or more videos into multiple steps" selected in each of areas 111 to 116, and it is possible to transition from screen 110 to the next screen. Note that temporary generation area 117 may be in a state where it cannot be selected until selections in areas 111 to 116 are completed.

[0039] 1B, an example of selecting the "step granularity" in area 111 has been described, but the present invention is not limited to this. For example, in area 111, it may be possible to select the number of steps in the electronic manual (e.g., one of 2, 3, 4, 5, 6, 7, 8, 9, and 10).

[0040] 1C shows an example of a screen 120 displayed on a user device. The screen 120 is a preview screen for viewing a provisionally generated electronic manual. The screen 120 is an example of a screen transitioned to from the screen 110 shown in FIG. 1B when the provisional generation area 117 in the screen 110 shown in FIG. 1B is selected by the user.

[0041] 1C, the screen 120 includes an outline area 121 for providing an outline of the provisionally generated electronic manual, a step area 122 for displaying each of a plurality of steps, a redo area 123 for redoing the provisional generation of the electronic manual, an editing start area 124 for executing editing of the provisionally generated electronic manual, and a final generation area 125 for executing the final generation of the provisionally generated electronic manual. The redo area 123, the editing start area 124, and the final generation area 125 are configured to be selectable.

[0042] The summary of the provisionally generated electronic manual displayed in the summary area 121 may be input before the provisional generation of the electronic manual, or may be automatically generated during the provisional generation of the electronic manual. In the example shown in FIG. 1C, the screen 120 displays a first step, a second step, and a part of a third step among the multiple steps. However, the user can view all of the multiple steps by performing a predetermined operation (e.g., vertical scrolling). When the user selects the redo area 123, the screen 120 transitions to the screen 110 of FIG. 1B, where the user can redo the input of the conditions for converting audio included in one or more videos into multiple steps. When the user selects the editing start area 124, the screen 120 transitions to the screen 130 of FIG. 1D, where the user can edit the provisionally generated electronic manual. When the user selects the final generation area 125, the final generation of the provisionally generated electronic manual is executed.

[0043] 1C, each step area 122 includes an image area 126 for displaying a sub-movie or a still image divided from one or more movies selected in the movie selection area 101 of FIG. 1A, a title area 127 for displaying the title of the step, and an explanation area 128 for displaying an explanation of the step. The image area 126 may display a movie, as in the image area of ​​the first step, or a still image, as in the image area of ​​the second step. When a movie is displayed in the image area 126, the image area 126 is configured to be selectable, and the movie can be played in response to a user operation (e.g., tapping, clicking, or hovering) for selecting the image area 126.

[0044] The number of steps displayed on screen 120, the number of characters in the title of each step, the number of characters in the explanation for each step, and the wording of the explanation for each step are in accordance with the "conditions for converting audio contained in one or more videos into multiple steps" selected in each of areas 111 to 114 on screen 110 of FIG. 1B. Furthermore, the presence or absence of subtitles in the video displayed in image area 126 for each step is in accordance with the "conditions for converting audio contained in one or more videos into multiple steps" selected in area 116 on screen 110 of FIG. 1B. The audio contained in one or more videos may be in a colloquial style, while the title and explanation for each step may be in a formal style.

[0045] 1D shows an example of a screen 130 displayed on a user device. The screen 130 is a screen for editing a provisionally generated electronic manual. The screen 130 is an example of a screen transitioned to from the screen 120 shown in FIG. 1C when the user selects the editing start area 124 in the screen 120 shown in FIG. 1C.

[0046] 1D, screen 130 includes video area 131 for displaying a video, step sequence area 132 for displaying a sequence of multiple steps, indicator area 133 for displaying an indicator for editing one or more videos, division area 134 for dividing one or more videos, and end editing area 135 for ending editing of the provisionally generated electronic manual. Dividing area 134 and end editing area 135 are configured to be selectable.

[0047] Indicator area 133 horizontally represents a timeline of one or more videos selected in video selection area 101 of Fig. 1A. The left end of indicator area 133 may be, for example, the playback start time (i.e., 0 minutes 0 seconds) of one or more videos selected in video selection area 101 of Fig. 1A, and the right end of indicator area 133 may be, for example, the playback end time (e.g., M minutes S seconds) of one or more videos selected in video selection area 101 of Fig. 1A. Here, M is an integer from 0 to 59, and S is an integer from 1 to 59.

[0048] In the example shown in FIG. 1D, the indicator area 133 includes a current position indicator 136 indicating the current playback position, a division position indicator 137 indicating the division position of the video that was automatically divided when the provisional generation of the electronic manual was performed, a candidate division time period indicator 138 indicating a candidate division time period between steps of the provisionally generated electronic manual, and a deletion time period indicator 139 indicating a time period of the video that was automatically deleted for a specified reason (e.g., no change in the image appears for a specified time period) when the provisional generation of the electronic manual was performed.

[0049] A video at a playback time corresponding to the location of current position indicator 136 is displayed in video area 131. Current position indicator 136 can slide horizontally on indicator area 133. When a user selects split area 134, a split position indicator 137 can be placed at the position of indicator area 133, making it possible to split one or more videos at the position of indicator area 133.

[0050] The displayed division position indicator 137 may be, for example, slidable horizontally within the division candidate time period indicator 138, thereby making it possible to adjust the division position within the division candidate time period between steps of the provisionally generated electronic manual. Note that the displayed division position indicator 137 may also be slidable horizontally beyond the division candidate time period indicator 138.

[0051] In the example shown in FIG. 1D , the step sequence area 132 includes join indicators 140 for joining adjacent steps. The number of join indicators 140 corresponds to the number of split position indicators 137 displayed in the indicator area 133. The join indicators 140 are configured to be selectable. By selecting a join indicator 140, the user can remove the selected join indicator 140 and combine two adjacent steps into one step. At this time, the split position indicator 137 corresponding to the removed join indicator 140 also disappears.

[0052] The user can restore the automatically deleted video by performing a predetermined operation on the deletion time zone indicator 139.

[0053] 1A, the multiple videos may be displayed consecutively in indicator area 133. In this case, a sequence change area (not shown) for changing the sequence of the multiple videos may be displayed on screen 130, and the sequence of the multiple videos may be changed in indicator area 133 according to the user's selection of the sequence change area.

[0054] In this way, the user can easily provisionally and finally generate an electronic manual (for example, a step-structured electronic manual) by selecting one or more videos that will form the basis of the electronic manual and inputting "conditions for converting audio contained in one or more videos into multiple steps." The user can also easily edit the provisionally generated electronic manual using the indicator area 133, current position indicator 136, division position indicator 137, division candidate time zone indicator 138, deletion time zone indicator 139, and combination indicator 140 as guides.

[0055] 2. System configuration for supporting the creation of electronic manuals FIG. 2 shows an example of the configuration of a system 200 for supporting the creation of an electronic manual.

[0056] In the embodiment shown in FIG. 2, the system 200 includes a computer system 210 for supporting the creation of an electronic manual, and user devices 2201 to 220. N The computer system 210 is provided with the user devices 2201 to 220 via the Internet 230. N The user devices 2201 to 220 are configured to be able to communicate with each of the user devices 2201 to 220. N can be manipulated by a user who wishes to create an electronic manual, where N is an integer greater than or equal to 1.

[0057] Computer system 210 is an information processing system that executes processes for a management company that provides and manages programs for supporting the creation of electronic manuals. In the embodiment shown in Fig. 2, computer system 210 includes an interface unit 211, a processor unit 212 including one or more CPUs (Central Processing Units), and a memory unit 213. The hardware configuration of computer system 210 is not particularly limited as long as it can realize its functions, and may be configured as a single machine or a combination of multiple machines.

[0058] The interface unit 211 is connected to the user devices 2201 to 220 N Controls communication with each of the

[0059] The memory unit 213 stores programs required to execute processing, data required to execute the programs, and the like. Here, it does not matter how the programs are stored in the memory unit 213. For example, the programs may be pre-installed in the memory unit 213. Alternatively, the programs may be installed in the memory unit 213 by being downloaded via a network such as the Internet 230, or may be installed in the memory unit 213 via a storage medium such as an optical disk or USB.

[0060] The processor unit 212 controls the overall operation of the computer system 210. The processor unit 212 reads and executes programs stored in the memory unit 213. This allows the computer system 210 to function as a device that executes desired steps, and the processor unit 212 of the computer system 210 to operate as a means for achieving desired functions.

[0061] 2, the computer system 210 is connected to a database unit 240. The database unit 240 can store, for example, an electronic manual that has been temporarily generated and then fully generated.

[0062] The user device 2201 is configured to be able to communicate with the computer system 210 via the Internet 230. In the embodiment shown in FIG. 2, the user device 2201 includes an interface unit 221, a processor unit 222, a memory unit 223, a display unit 224, and an input unit 225 for receiving input (e.g., sound, input by selection (e.g., tap, click)). The user device 2201 may further include, for example, an output unit (not shown) for outputting output (e.g., sound). The user device 2201 may be a mobile wireless terminal such as a mobile phone, smartphone, or tablet terminal, or a personal computer such as a laptop PC or notebook PC. The configurations of the interface unit 221, processor unit 222, and memory unit 223 of the user device 2201 are similar to those of the interface unit 211, processor unit 212, and memory unit 213 of the computer system 210, and therefore detailed description thereof will be omitted here. The memory unit 223 stores one or more videos that can serve as the basis for an electronic manual. User devices 2202 to 220 N The same is true for .

[0063] In the embodiment shown in FIG. 2, the user devices 2201 to 220 N Although each of these is described as being capable of communicating with computer system 210 via Internet 230, the invention is not so limited. Any type of network may be used in place of Internet 230.

[0064] 2, the database unit 240 is provided outside the computer system 210, but the present invention is not limited to this. The database unit 240 can also be provided inside the computer system 210. The configuration of the database unit 240 is not limited to a specific hardware configuration. For example, the database unit 240 may be configured as a single hardware component, or may be configured as multiple hardware components. For example, the database unit 240 may be configured as a single external hard disk drive of the computer system 210, or may be configured as cloud storage connected via a network.

[0065] 3. Processing performed in a computer system Fig. 3 shows an example of processing executed in computer system 210. Each step shown in Fig. 3 is executed, for example, by processor unit 212 of computer system 210. Each step shown in Fig. 3 will be described below.

[0066] Step S301: One or more videos are identified. The computer system 210 receives one or more videos, for example, from the user device 2201, and is thereby able to identify the one or more videos. The identified one or more videos are, for example, videos that a user operating the user device 2201 wants to use as the basis for an electronic manual. This process may correspond to, for example, an operation on the video selection area 101 in FIG. 1A.

[0067] At this time, the computer system 210 may receive input for setting the input language (i.e., the language of the audio included in one or more videos) and the output language (i.e., the language of the electronic manual in the provisional generation and final generation). This process may correspond to, for example, operations on the input language setting area 102 and the output language setting area 103 in FIG. 1A. By receiving an input for setting the input language, the computer system 210 can improve the accuracy of the structured text in step S308. Furthermore, by receiving an input for setting the output language, the computer system 210 can create an electronic manual in either the same language as the input language or a language different from the input language.

[0068] Step S302: It is determined whether the one or more videos received in step S301 contain audio. The audio contained in the one or more videos may be audio indicating the procedures of an electronic manual. If the determination result is "Yes", the process proceeds to step S307; if the determination result is "No", the process proceeds to step S303.

[0069] Step S303: A process is executed to warn that the one or more videos received in step S301 do not contain audio. This process may be achieved, for example, by the computer system 210 transmitting a warning to the user device 2201 indicating that the one or more videos do not contain audio and presenting the warning on the user device 2201, or by the computer system 210 transmitting a warning sound signal to the user device 2201 indicating that the one or more videos do not contain audio and emitting the warning sound on the user device 2201.

[0070] Step S304: It is determined whether a user input indicating that speech is to be input is received. The user input indicating that speech is to be input may be received, for example, from the user device 2201. If the determination result is "Yes", the process proceeds to step S306, and if the determination result is "No", the process proceeds to step S305.

[0071] Step S305: A process is executed to notify the user that the electronic manual cannot be created. This process may be achieved, for example, by the computer system 210 transmitting information indicating that the electronic manual cannot be created to the user device 2201 and presenting the information on the user device 2201.

[0072] Step S306: It is determined whether audio input has been received. The audio input may be received, for example, from the user device 2201. The audio input may be achieved, for example, by inputting pre-recorded audio, or by recording audio in parallel with playing one or more videos on the user device 2201. If the determination result is "Yes," the process proceeds to step S307; if the determination result is "No," the process returns to step S306.

[0073] Step S307: Conditions for converting audio included in one or more videos into a plurality of steps are identified. The conditions for converting into a plurality of steps include at least a limit on the number of steps, which may correspond to the operation on region 111 in FIG. 1B. The conditions for converting into a plurality of steps may further include a limit on the number of characters in the title (e.g., a limit on the number of characters in the title of each step in the electronic manual) and / or a limit on the number of characters in the description (e.g., a limit on the number of characters in the description of each step in the electronic manual), which may correspond to the operation on regions 112 and 113 in FIG. 1B. The conditions for converting into a plurality of steps may further include a limit on the wording of the description (e.g., a limit on the wording of the description of each step in the electronic manual) and / or an expected reader of the electronic manual, which may correspond to the operation on regions 114 and 115 in FIG. 1B.

[0074] Step S308: Structured text for constituting the multiple steps of the electronic manual is generated. The structured text is generated from the audio included in the one or more videos identified in step S301 based on conditions for converting the audio included in the one or more videos into multiple steps. The structured text includes at least a title or description for each of the multiple steps. The title of each of the multiple steps included in the structured text may correspond, for example, to the description in title area 127 on screen 120 of FIG. 1C. The description of each of the multiple steps included in the structured text may correspond, for example, to the description in description area 128 on screen 120 of FIG. 1C. The structured text may be generated using, for example, artificial intelligence (e.g., ChatGPT). The computer system 210 is configured to convert the structured text (particularly, the title or description for each of the multiple steps included in the structured text) from an input language to an output language. This allows the computer system 210 to generate the structured text in the set output language even if the input language of the audio included in one or more videos is different from the output language of the electronic manual.

[0075] Note that computer system 210 may generate the structured text directly from the audio included in one or more videos based on the conditions for converting the audio included in one or more videos into a plurality of steps. Alternatively, computer system 210 may convert the audio included in one or more videos into text by transcribing the audio included in one or more videos, and generate the structured text based on the converted text and the conditions for converting the audio included in one or more videos into a plurality of steps.

[0076] Step S309: One or more videos are divided into multiple sub-videos or still images. This process is performed based at least on the one or more videos identified in step S301 and the structured text generated in step S308. This process can be achieved by the computer system 210, for example, identifying scene change timings within the videos based on the structured text and generating multiple sub-videos or still images by dividing the one or more videos based on the scene change timings. The identification of the scene change timings can be achieved, for example, by identifying breaks in the content of the structured text based on the structured text and identifying timings in the audio corresponding to the breaks in the structured text as the scene change timings. The breaks in the content of the structured text can exist, for example, between steps of multiple steps.

[0077] The computer system 210 may generate multiple sub-movies or still images from one or more videos by, for example, identifying timings of significant image changes in one or more videos, identifying timings of audio breaks, and dividing one or more videos at timings where the timings of significant image changes, scene changes, and audio breaks coincide. The timings of significant image changes in one or more videos may be, for example, timings where the area of ​​the image that has changed relative to the display area of ​​the video exceeds a predetermined threshold. The timings of audio breaks may be, for example, timings where a period of silence in one or more videos exists for a predetermined length of time.

[0078] The computer system 210 may accomplish segmenting one or more videos into multiple sub-videos or still images by, for example, segmenting one or more videos into multiple candidate sub-videos based at least on the one or more videos and the structured text, identifying a candidate sub-video among the multiple candidate sub-videos that has audio above a predetermined volume for a predetermined time while showing no image changes, and converting the candidate sub-video into a still image based on the candidate sub-video. Converting a candidate sub-video into a still image may be accomplished, for example, by capturing a portion of the candidate sub-video as a still image.

[0079] Step S310: A provisional electronic manual is generated. This process is performed based on the structured text generated in step S308 and the multiple sub-videos or still images generated in step S309. This process may be performed in response to receiving a user input from the user device 2201 requesting provisional generation of an electronic manual. This process may correspond to, for example, an operation on the provisional generation area 117 in FIG. 1B. The provisionally generated electronic manual may not include audio contained in one or more videos. The language of the provisionally generated electronic manual may be changed from the input language to the output language according to the language settings in the input language setting area 102 and the output language setting area 103 in FIG. 1A. If the input language of the audio contained in one or more videos is different from the output language of the electronic manual, the language of the provisionally generated electronic manual may be changed, for example, by machine translation. In addition, when the structured text (especially the titles or descriptions of each of the multiple steps included in the structured text) has been converted into an output language, the computer system 210 can provisionally generate an electronic manual based on the titles or descriptions of each of the multiple steps converted into the output language and the multiple sub-videos or still images generated in step S309.

[0080] Step S311: It is determined whether a user input for executing the book generation of the electronic manual has been received. The user input for executing the book generation of the electronic manual may be received, for example, from the user device 2201. If the determination result is "Yes", the process proceeds to step S312, and if the determination result is "No", the process returns to step S311.

[0081] Step S312: Final generation of the electronic manual is executed. This completes the electronic manual. The final generated electronic manual may be output according to the language setting in the output language setting area 103 of FIG. 1A. When final generation of the electronic manual is executed, the computer system 210 may generate audio data for reading aloud the structured text generated in step S308. This makes it possible to realize automatic reading of the completed electronic manual. Furthermore, by generating audio data in multiple languages ​​for reading aloud the structured text generated in step S308, it is possible to provide the electronic manual in multiple languages, regardless of the language of the audio contained in one or more videos.

[0082] Fig. 4 shows another example of processing executed in the computer system 210. The steps shown in Fig. 4 are executed, for example, by the processor unit 212 of the computer system 210. The steps shown in Fig. 4 show an example of processing for editing a provisionally generated electronic manual at any timing after step S311 and before step S312 in Fig. 3. Each step shown in Fig. 4 will be described below.

[0083] Step S401: A user input indicating a desire to edit the provisionally generated electronic manual is received. The user input indicating a desire to edit the provisionally generated electronic manual may be received, for example, from the user device 2201. This process may correspond to, for example, an operation on the edit start area 124 in FIG. 1C.

[0084] Step S402: A division candidate time period between the steps of the provisionally generated electronic manual is identified. Within the division candidate time period, the user may be able to adjust the division position between the steps of the provisionally generated electronic manual. The division candidate time period is identified, for example, based on the structured text generated in step S308 and the audio included in one or more videos. Specifically, the computer system 210 may identify the playback time of the audio corresponding to each step based on the structured text generated in step S308, and identify the division candidate time period based on the playback time of the audio corresponding to each step. For example, the division candidate time period may be the entire time period between the end of the playback time of the audio corresponding to a step and the start of the playback time of the audio corresponding to the next step, or may be a time period within a predetermined range from a certain point between the end of the playback time of the audio corresponding to a step and the start of the playback time of the audio corresponding to the next step.

[0085] Step S403: A process is executed to present time periods of division candidates between the steps of the plurality of steps in the provisionally generated electronic manual. This process may be achieved, for example, by the computer system 210 transmitting information indicating time periods of division candidates between the steps of the provisionally generated electronic manual to the user device 2201 and presenting the information on the user device 2201. This process may correspond to, for example, displaying the image 130 in FIG. 1D on the user device 2201. This allows the user to start editing the provisionally generated electronic manual.

[0086] Step S404: It is determined whether a user input for editing the provisionally generated electronic manual has been received. The user input for editing the provisionally generated electronic manual may be received, for example, from the user device 2201. This process may correspond to, for example, an operation on the image 130 in FIG. 1D. If the determination result is "Yes," the process proceeds to step S405; if the determination result is "No," the process proceeds to step S406.

[0087] The user input for editing the provisionally generated electronic manual may include, for example, a user input for adjusting a division position between steps of the provisionally generated electronic manual within a time period of a division candidate. This may correspond to an operation on the division position indicator 137 in FIG. 1D . The user input for editing the provisionally generated electronic manual may also include, for example, a user input for dividing one or more videos. This may correspond to an operation on the division area 134 in FIG. 1D . The user input for editing the provisionally generated electronic manual may also include a user input for restoring automatically deleted videos. This may correspond to an operation on the deletion time period indicator 139 in FIG. 1D . The user input for editing the provisionally generated electronic manual may also include a user input for merging adjacent steps. This may correspond to an operation on the merge indicator 140 in FIG. 1D .

[0088] Step S405: Editing of the temporarily generated electronic manual is executed in response to a user input for editing the temporarily generated electronic manual. For example, if the user input for editing the temporarily generated electronic manual is a user input for adjusting the division positions between the steps of the temporarily generated electronic manual within the time period of the division candidate, adjustment of the division positions between the steps of the temporarily generated electronic manual within the time period of the division candidate is executed.

[0089] Step S406: It is determined whether a user input to end editing of the provisionally generated electronic manual has been received. The user input to end editing of the provisionally generated electronic manual may be received, for example, from the user device 2201. This process may correspond to, for example, an operation on the editing end area 135 in FIG. 1D. [Example]

[0090] An example of generating structured text using ChatGPT will be described below.

[0091] For example, suppose the audio included in one or more videos is: "First, open the Settings app. After opening the Settings app, open General management at the bottom of the left menu. Then, tap Text-to-Speech in the right menu. Next, tap the gear icon next to Priority Engine. The installed languages ​​are listed at the bottom right. If your language is not listed, tap Install Audio Data. If you tap a language that is not installed, a screen like this will appear prompting you to install it. Press the Install button. When the installation is complete, a screen like this will appear notifying you. This completes the audio data installation process." Using ChatGPT under the following conditions: "The number of steps is limited to 50," "The title is limited to 50 characters," and "The description is limited to 200 characters," the following output sentence can be output as structured text from this audio: [ka]

[0092] In the above output sentence, "Step 0:..." represents the title of each step, and "Description" represents the description of each step. Also, "Base Text" refers to the text that is generated from the audio corresponding to each step. In the above audio example, the structured text includes seven steps.

[0093] When generating the structured text, a correspondence between each step of the multiple steps and the audio duration corresponding to each step is maintained and / or recorded. In the audio example described above, the base text of step 1, "First, open the Settings app.", corresponds to the audio duration from 0 minutes 0 seconds to 0 minutes 18 seconds in one or more videos; the base text of step 2, "Open General management at the bottom of the left menu.", corresponds to the audio duration from 0 minutes 19 seconds to 0 minutes 23 seconds; the base text of step 3, "Tap Read Text to Speech.", corresponds to the audio duration from 0 minutes 24 seconds to 0 minutes 30 seconds; the base text of step 4, "Tap the gear icon next to the priority engine.", corresponds to the audio duration from 0 minutes 32 seconds to 0 minutes 38 seconds; and the base text of step 5, "Already installed, The installed languages ​​are listed at the bottom right, and if your language is not listed, tap Install audio data. "Tap a language that is not installed from among these" corresponds to the audio playback time from 0 minutes 40 seconds to 0 minutes 48 seconds, the base text of step 6 "When you tap a language that is not installed, a screen like this will appear prompting you to install it, so press the Install button" corresponds to the audio playback time from 0 minutes 50 seconds to 1 minute 02 seconds, and the base text of step 7 "When the installation is complete, a screen like this will appear notifying you of the completion" corresponds to the audio playback time from 1 minute 04 seconds to 1 minute 15 seconds. In this case, for example, since step 1 of the structured text corresponds to the audio playback time of "0 minutes 0 seconds to 0 minutes 18 seconds," and step 2 of the structured text corresponds to the audio playback time of "0 minutes 19 seconds to 0 minutes 23 seconds," computer system 210 can identify the candidate division time period between step 1 and step 2 as "0 minutes 18 seconds to 0 minutes 19 seconds," which is between "0 minutes 0 seconds to 0 minutes 18 seconds" and "0 minutes 19 seconds to 0 minutes 23 seconds."When performing provisional generation of an electronic manual, it is possible to determine the division position of the video between step 1 and step 2 within the candidate division time period of "0 minutes 18 seconds to 0 minutes 19 seconds."The same applies to the time slots of the split candidates between steps 2 and 3, between steps 3 and 4, between steps 4 and 5, between steps 5 and 6, and between steps 6 and 7.

[0094] When the computer system 210 determines the division position of the video between step 1 and step 2 within the division candidate time period "0:18 to 0:19", it may, for example, automatically determine the division position to be in the center of the division candidate time period "0:18 to 0:19", or may determine the division position randomly within the division candidate time period "0:18 to 0:19". The same applies to the other division candidate time periods.

[0095] 3 and 4, an example has been described in which computer system 210 executes the processing of each step shown in FIGS. 3 and 4, but the present invention is not limited to this. For example, the processing of each step shown in FIGS. 3 and 4 may be executed by, for example, user device 2201 (particularly, processor unit 222 of user device 2201) instead of computer system 210. In this case, in step S301 of FIG. 3, one or more videos among multiple videos stored in memory unit 223 of user device 2201 are selected in video selection area 101 of FIG. 1A, thereby enabling user device 2201 to identify one or more videos that will serve as the basis for the electronic manual. In step S307 of FIG. 3, a "condition for converting audio included in one or more videos into multiple steps" is selected in each of areas 111 to 116 of FIG. 1B, thereby enabling user device 2201 to identify the condition for converting audio included in one or more videos into multiple steps. 4. Furthermore, the user device 2201 receives user input via the input unit 225 of the user device 2201 in steps S304 and S311 of FIG. 3 and in step S406 of FIG.

[0096] In the embodiment shown in Figures 3 and 4, an example has been described in which the processing of each step shown in Figures 3 and 4 is realized by the processor executing a program stored in the memory, but the present invention is not limited to this. At least some of the processing of each step shown in Figures 3 and 4 may be realized by a hardware configuration such as a control circuit.

[0097] As described above, the present invention has been illustrated using the preferred embodiment of the present invention, but the present invention should not be interpreted as being limited to this embodiment. It is understood that the scope of the present invention should be interpreted only by the claims. It is understood that a person skilled in the art can implement an equivalent scope based on the description of the present invention and common technical knowledge from the description of the specific preferred embodiment of the present invention. [Industrial Applicability]

[0098] The present invention is useful for reducing the time and effort required to create an electronic manual by providing a computer system, a program, etc. for supporting the creation of an electronic manual. [Explanation of symbols]

[0099] 200 systems 210 Computer Systems 2201~220 N User Device 230 Internet 240 Database Department

Claims

1. A computer system for supporting the creation of an electronic manual, the computer system comprising: means for receiving one or more videos; means for receiving information indicating a condition for converting to a plurality of steps; means for generating structured text for composing a plurality of steps from audio included in the one or more videos based on the conditions, the structured text including at least a title or description of each of the plurality of steps; means for dividing the one or more videos into a plurality of sub-videos or still images based at least on the one or more videos and the structured text; means for provisionally generating the electronic manual based on the structured text and the plurality of sub-moving images or still images; A computer system comprising:

2. The computer system of claim 1 , wherein the condition includes a limit on the number of steps.

3. The computer system of claim 2 , wherein the conditions further include a limit on the number of characters in a title and / or a limit on the number of characters in a description.

4. The computer system according to claim 1 , wherein the audio included in the one or more videos is audio that indicates a procedure in the electronic manual.

5. The computer system according to claim 4 , wherein the provisionally generated electronic manual does not include audio included in the one or more videos.

6. Dividing the one or more videos into a plurality of sub-videos or still images based at least on the one or more videos and the structured text includes: Segmenting the one or more videos into a plurality of candidate sub-videos based at least on the one or more videos and the structured text; Among the plurality of candidate secondary videos, a candidate secondary video in which a sound exceeding a predetermined volume is present for a predetermined time but no change in image is displayed is converted into a still image based on the candidate secondary video.

10. The computer system of claim 1, comprising:

7. Dividing the one or more videos into a plurality of sub-videos or still images based at least on the one or more videos and the structured text includes: Identifying timing of scene changes based on the structured text; generating the plurality of sub-moving images or still images by dividing the one or more moving images based on the timing of the scene changes; 10. The computer system of claim 1, comprising:

8. Identifying the timing of the scene change based on the structured text includes: Identifying breaks in the content of the structured text based on the structured text; identifying a timing in the audio corresponding to a break in the structured text as a timing of a change in the scene; 8. The computer system of claim 7, comprising:

9. Dividing the one or more videos into a plurality of sub-videos or still images based at least on the one or more videos and the structured text includes: Identifying timings of large image changes in the one or more moving images; Identifying timing of the audio break; generating the plurality of sub-moving images or still images by dividing the one or more moving images at timings where the timing of the large image change, the timing of the scene change, and the timing of the audio break coincide; The computer system of claim 7 further comprising:

10. Generating the structured text from the audio included in the one or more videos based on the condition includes: converting the audio contained in the one or more videos into text by transcribing the audio; generating the structured text based on the text converted from the speech and the conditions; 10. The computer system of claim 1, comprising:

11. The computer system includes: means for receiving a first user input indicating a desire to edit the provisionally generated electronic manual; a means for specifying a division candidate time period between steps of the provisionally generated electronic manual in response to receiving the first user input, wherein a user can adjust a division position between steps of the provisionally generated electronic manual within the division candidate time period; means for presenting the division candidate time slots; means for receiving a second user input for adjusting a division position between steps of the provisionally generated electronic manual within the time period of the division candidate; means for editing the provisionally generated electronic manual based on the second user input; The computer system of claim 1 further comprising:

12. The computer system of claim 11 , wherein identifying the candidate segmentation time periods includes identifying the candidate segmentation time periods based on the structured text and audio included in the one or more videos.

13. Identifying the segmentation candidate time slots based on the structured text and audio included in the one or more videos includes: Identifying a playback time of the audio corresponding to each step of the plurality of steps based on the structured text; Identifying the time slots of the division candidates based on the playback times of the audio corresponding to each step; 12. The computer system of claim 11, comprising:

14. The computer system includes: means for receiving a third user input for executing the book generation of the electronic manual; means for executing book generation of the electronic manual in response to receiving the third user input; The computer system of claim 1 further comprising:

15. The computer system includes: means for determining whether the one or more videos contain audio; means for warning a user that the one or more videos do not contain audio when the one or more videos are determined not to contain audio; The computer system of claim 1 further comprising:

16. The computer system of claim 1 , wherein the audio included in the one or more videos is in a colloquial style, and the title and the description are in a written style.

17. The computer system of claim 1 , further comprising means for generating audio data for reading the structured text aloud.

18. The computer system includes: means for receiving input for setting an input language and an output language; means for converting the language of the title or the description of each of the plurality of steps included in the structured text from the input language to the output language; Equipped with The computer system of claim 1, wherein provisionally generating the electronic manual based on the structured text and the plurality of sub-videos or still images includes provisionally generating the electronic manual based on the title or the description converted into the output language of each of the plurality of steps and the plurality of sub-videos or still images.

19. A program executed in a computer system for supporting the creation of an electronic manual, the computer system comprising: a processor unit for controlling the operation of the computer system; When the program is executed by the processor unit, receiving one or more videos; receiving information indicating a condition for converting to a plurality of steps; generating structured text for composing a plurality of steps from audio included in the one or more videos based on the conditions, the structured text including at least a title or description of each of the plurality of steps; Segmenting the one or more videos into a plurality of sub-videos or still images based at least on the one or more videos and the structured text; provisionally generating the electronic manual based on the structured text and the plurality of sub-moving images or still images; a program causing the processor unit to perform at least the above.

20. A program for assisting in the creation of an electronic manual, the program being executed on a user device, the user device having a processor unit for controlling an operation of the user device; When the program is executed by the processor unit, Identifying one or more videos; identifying information indicating a condition for converting to a plurality of steps; generating structured text for composing a plurality of steps from audio included in the one or more videos based on the conditions, the structured text including at least a title or description of each of the plurality of steps; Segmenting the one or more videos into a plurality of sub-videos or still images based at least on the one or more videos and the structured text; provisionally generating the electronic manual based on the structured text and the plurality of sub-moving images or still images; a program causing the processor unit to perform at least the above.

Citation Information

Patent Citations

  • Image layout apparatus. image layout program, and image layout method

    JP2004120127A

  • Program, method and device for supporting creation of document

    JP2019020789A

  • Explicit knowledge formalization system and method of the same

    JP2019144822A

  • Video manual creation device, video manual creation method, and video manual creation program

    JP7023427B1

  • Program, server, and system for managing work logs on basis of electronic manual

    WO2017183064A1