Method and System for Automated Generation of Product Instruction Videos Using Learning-Based Standard Action Extraction
Patent Information
- Application Number
- KR1020250140615
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-09-29
- Publication Date
- 2026-08-14
- Estimated Expiration
- 2045-09-29
Smart Images

Figure 112025110678217-PAT00002_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to an artificial intelligence-based automatic content generation technology, and more specifically, to a method and system for automatically generating product description videos by learning product manuals from multiple industry groups to automatically extract common patterns (standard movements), and determining and applying animation parameters based on the extracted standard movements. Background Technology
[0002] Product manuals for various electronic products, furniture, automobiles, medical devices, etc., have provided essential procedures ranging from installation, assembly, initial setup, daily use, and maintenance, primarily through text and illustrations.
[0003] However, in actual user environments, manuals are rarely read thoroughly from beginning to end, and problems frequently arise where complex procedures or safety precautions are omitted or misunderstood.
[0004] Accordingly, the demand for instructional videos to improve product understanding and reduce installation and usage errors is continuously increasing.
[0005] However, conventional technology related to this has the following limitations.
[0006] Conventional technology involves manual video production methods that require the repetition of scenario writing, filming or 3D modeling, editing, and multilingual localization for each product and model, resulting in excessive costs and time, and large variations in quality depending on the producer.
[0007] Furthermore, many of the techniques claiming to be automation in conventional technology remain based on human-designed templates or rule-based assembly, making it difficult to generalize standard operations that appear commonly across various industries and product types from data. As a result, context adaptation is limited, such as adjusting explanation density based on the sequence of procedures, dependencies, or user proficiency.
[0008] Furthermore, manuals in conventional technology have limitations in that they differ in language, region, and expression style, and their formats, such as text, images, and tables, are heterogeneous, making it difficult to accept them as consistent input to automatically identify the step structure and visualize them with optimal timing, speed, and emphasis methods. Prior art literature
[0009] Korean Patent Publication No. 10 - 2801082 (Registration Date: April 22, 2025) The problem to be solved
[0010] The present invention aims to solve the aforementioned problems. One embodiment of the present invention provides a technology capable of automatically extracting common, repetitive actions across various industries and product types as standard actions by learning a large volume of product manuals based on data, and automatically generating product description videos based thereon.
[0011] One embodiment of the present invention integrates and learns multiple manuals collected from electronic products, furniture, automobiles, medical devices, etc., into a standard format, derives candidate patterns through action verb frequency, semantic clustering, and precedence dependency pattern learning, defines standard actions through statistical verification, and statistically derives the optimal timing, playback speed, step decomposition, and emphasis method for each standard action from the training data to determine and apply them as animation parameters, thereby providing an optimal completed video by automatically selecting and combining optimal clips through adaptive pattern matching with the step structure of the target manual.
[0012] One embodiment of the present invention enables significant cost and time reduction compared to manual production through the above process, while achieving consistency in visual style and standardization of explanatory quality, and providing customized multiple versions of clips that reflect context, such as product type, user level, and procedural complexity. Additionally, by applying transition pattern learning and playback timing control to perform natural connections between clips and eliminate duplication, the completeness of the final video is enhanced.
[0013] One embodiment of the present invention enables the provision of explanatory videos that reflect linguistic and regional differences through multilingual manual processing and automatic localization, and allows for the continuous improvement of standard operation definitions and clip quality through self-evolving updates upon the accumulation of new manuals and user feedback. This provides technical effects such as objective standardization, scalability, consistent quality, and adaptive customization.
[0014] However, the problems to be solved in this disclosure are not limited to those mentioned above, and may be expanded in various ways without departing from the spirit and scope of this disclosure. means of solving the problem
[0015] One technical aspect of the present invention proposes a method for automatically generating product description videos using learning-based standard motion automatic extraction. The method is performed in an automatic video generation device and comprises: a step of constructing a learning dataset based on a plurality of product manuals collected for a plurality of industry groups; a step of extracting common patterns that are commonly repeated in the plurality of product manuals from the learning dataset and defining standard motions based on statistical verification of the common patterns; a step of determining animation parameters for the standard motions and generating animation clips based on the animation parameters; and a step of adaptively calculating a correspondence relationship between a target manual and a standard motion, and generating a product description video by combining animation clips based on the calculated correspondence relationship.
[0016] In one embodiment, the step of constructing the training dataset may include: collecting the plurality of product manuals from the plurality of industry groups; and converting the plurality of product manuals into a structured standard format in procedural units to perform format integration.
[0017] In one embodiment, the step of defining the standard operation may include: extracting action verbs from a product manual and performing semantic clustering on the extracted action verbs to generate candidate patterns; and defining, among the generated candidate patterns, a candidate pattern that satisfies a safety verification criterion while satisfying a preset frequency threshold as the standard operation.
[0018] In one embodiment, the step of defining the standard operation may further include: a step of learning a pattern of sequential dependencies between operations in a product manual to calculate the dependency between the immediate and subsequent steps for the candidate pattern; and a step of defining the candidate pattern as a standard operation if the dependency is greater than or equal to a preset threshold value.
[0019] In one embodiment, the animation parameters include a camera viewpoint, playback speed, step decomposition level, and emphasis method. The step of determining the animation parameters may include: selecting a camera viewpoint for the standard motion from a set of candidate viewpoints based on the statistical derivation result of the training data; setting a playback speed of the standard motion based on the statistical value of the training data; determining a step decomposition level of the standard motion based on the level of granularity derived from the training data; and selecting an emphasis method of the standard motion based on the preference derived from the training data.
[0020] In one embodiment, the step of generating the animation clip may further include: a step of automatically generating the animation clip by simulating the standard motion in a 3D environment with a physics engine applied by applying the determined animation parameters; and a step of correcting the visual parameters to maintain consistency in visual style between the generated animation clips.
[0021] In one embodiment, the step of generating the animation clip may further include: generating the animation clip in a plurality of versions; assigning product type and parameter values as metadata to each version of the animation clip; and storing the plurality of versions of the animation clip and the metadata in a clip library.
[0022] In one embodiment, the step of generating a product description video by combining the animation clips may include: identifying step boundaries from the procedure description of a target manual and extracting key action candidates for each step; calculating a similarity index for each step based on semantic similarity, order suitability, and context suitability between the key action candidates and the standard action; and selecting the animation clip with the maximum similarity index as the optimal clip for the corresponding step.
[0023] Another technical aspect of the present invention proposes a system for automatically generating product description videos using learning-based standard motion automatic extraction. The system comprises: a user terminal providing a target manual; and a video automatic generation server that receives the target manual from the user terminal, automatically generates a product description video based on the target manual, and provides it to the user terminal. The video automatic generation server constructs a learning dataset based on a plurality of product manuals collected for a plurality of industry groups, extracts common patterns that are commonly repeated in the plurality of product manuals from the learning dataset, defines a standard motion based on statistical verification thereof, determines animation parameters for the standard motion, generates an animation clip based on the animation parameters, adaptively calculates the correspondence relationship between the target manual and the standard motion, and generates the product description video by combining the animation clip based on the correspondence relationship. Effects of the invention
[0024] According to various embodiments of the present invention, by learning and verifying commonly repeated patterns from a large volume of product manuals to define standard operations and directly applying them as animation parameters, the generalization limitations of conventional template / rule-based approaches can be overcome, and high-quality product description videos can be automatically generated. Accordingly, even for complex installation, assembly, and usage procedures, core operations can be accurately visualized to enhance user understanding.
[0025] According to various embodiments of the present invention, camera viewpoints, playback speeds, step breakdowns, and emphasis methods for each standard operation are determined based on data to ensure visual representation of consistent quality. Therefore, quality variations depending on the manufacturer or outsourcing company are minimized, and standardization of explanatory quality can be achieved even if the product or model changes.
[0026] According to various embodiments of the present invention, the step structure of a target manual is analyzed, and the correspondence relationship with standard operations is adaptively calculated to automatically select and combine optimal animation clips. Accordingly, the manual editing process is significantly reduced, thereby having the effect of substantially reducing the costs and time required for planning, filming, editing, and inspection.
[0027] According to various embodiments of the present invention, multiple versions of clips can be automatically generated and managed for the same standard operation according to product type, user level, and procedural complexity, thereby enabling the provision of context-customized videos for beginners, experts, etc. Accordingly, the effect of optimizing the difficulty of explanation according to user groups and usage situations can be achieved.
[0028] According to various embodiments of the present invention, by reusing generated animation clips by creating a library with metadata, rapid video combination is possible for new products or derivative models, and the creation of organizational-level knowledge assets is promoted. Accordingly, operational scalability and maintenance efficiency for large-scale product families are improved.
[0029] According to various embodiments of the present invention, as new manual data or user feedback accumulates, standard operations and parameters can be self-updated, so the quality of explanation continuously improves over time. Therefore, it has the effect of rapidly adapting to new products or new procedures.
[0030] The effects obtainable from the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which the present disclosure belongs from the description below. Brief explanation of the drawing
[0031] FIG. 1 is a block diagram of a product description video automatic generation system (10) according to one embodiment of the present invention. FIG. 2 is a flowchart illustrating a method for automatically generating product description videos using learning-based standard motion automatic extraction according to an embodiment of the present invention. FIG. 3 is a block diagram of a learning dataset construction pipeline within an image automatic generation device (200) according to one embodiment of the present invention. FIG. 4 is a flowchart illustrating a standard operation automatic definition procedure according to one embodiment of the present invention. FIG. 5 is a flowchart illustrating a procedure for determining animation parameters according to an embodiment of the present invention. FIG. 6 is a flowchart illustrating the animation clip creation and library management procedure according to one embodiment of the present invention. FIG. 7 is a flowchart illustrating an adaptive matching and animation clip selection procedure between a target manual and a standard operation according to an embodiment of the present invention. FIG. 8 is a flowchart illustrating the procedure for providing final synthesis and rendering of a product description video according to an embodiment of the present invention. Specific details for implementing the invention
[0032] Hereinafter, embodiments of the present disclosure are described in detail with reference to the drawings so that those skilled in the art can easily practice them. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein. In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and brevity.
[0033] The various embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, phrases such as "A or B," "at least one of A and B," "at least one of A or B," "A, B or C," "at least one of A, B and C," and "at least one of A, B, or C" may each include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used simply to distinguish said components from other said components and do not limit said components in any other aspect (e.g., importance or order). Where any (e.g., 1st) component is referred to as "coupled" or "connected" to another (e.g., 2nd) component, with or without the terms "functionally" or "communicationly," it means that said any component may be connected to said other component directly (e.g., via a wire), wirelessly, or through a third component.
[0034] As used in the various embodiments of this document, the term “module” may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions.
[0035] Various embodiments of this document may be implemented as software (e.g., a program) comprising one or more instructions stored in a storage medium (e.g., memory) readable by a machine or device. For example, the processor of the machine or device may call at least one of the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by a machine may be provided in the form of a non-transitory storage medium. Here, "non-transitory" simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and this term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily in the storage medium.
[0036] According to one embodiment, the method according to the various embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0037] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
[0038] In this disclosure, the term "processor" may refer to hardware capable of performing functions and operations according to each designation described herein, computer program code capable of performing specific functions and operations, or an electronic recording medium loaded with computer program code capable of performing specific functions and operations. According to an embodiment, the operation of the processor may be defined and / or interpreted as the operation of a knowledge graph adjustment device, but is not limited thereto. The term "processor" refers to a functional and / or structural combination of hardware for carrying out the technical concept of this disclosure and / or software for driving said hardware.
[0040] FIG. 1 is a block diagram of a product description video automatic generation system (10) according to one embodiment of the present invention.
[0041] Generally, in conventional technology, when producing explanatory videos based on product manuals, a manual pipeline was primarily used in which a person wrote the script and repeatedly performed filming, editing, and localization.
[0042] This conventional method had limitations, such as production costs and time increasing exponentially as the number of products and models increased, significant quality variations depending on the manufacturer, and difficulty in accurately reflecting the sequence and dependencies of procedures or differences in user proficiency.
[0043] Furthermore, some techniques claiming to be automated remained limited to combinations of templates and rules, posing a problem in that it was difficult to generalize or verify common operation patterns (standard operations) that repeatedly appear in manuals across various industries based on data.
[0044] As a result, it was difficult to quickly video the complex installation, assembly, and initial setup procedures with consistent quality.
[0045] In contrast, a product description video automatic generation system (hereinafter abbreviated as 'system') (10) according to one embodiment of the present invention can automatically generate a product description video by learning common repeating patterns from a large amount of manual data to define standard movements, determining animation parameters (camera viewpoint, playback speed, step breakdown, emphasis method) for each standard movement based on data, and then automatically selecting and combining animation clips through adaptive matching with a target manual. This enables rapid, consistent, and low-cost video production for various products, languages, and skill levels.
[0046] Specifically, the system (10) includes a user terminal (100) and an image automatic generation device (200).
[0047] The user terminal (100) is a device that provides a target manual and receives a generated product description video, and can be implemented as various electronic devices such as a personal computer, laptop, smartphone, and tablet.
[0048] The user can provide the manual of the target product via upload or designated method, and receive the explanation video generated from the video automatic generation device (200) in the form of streaming or a file.
[0049] The video automatic generation device (200) receives input from a user terminal (100), and performs building a learning dataset, defining standard operations through common pattern extraction and statistical verification, determining animation parameters, generating animation clips, calculating correspondence relationships with target manuals and combining clips, and generating and providing a final video.
[0050] For example, the automatic video generation device (200) can perform the function of structuring and storing manuals collected from various industries into a standard format. At this time, different formats such as documents, images, tables, and captions are normalized in procedural units to preprocess them into a form suitable for subsequent learning.
[0051] For example, an automatic video generation device (200) can extract action verbs from the procedural text of a structured manual, generate and select candidate patterns by applying semantic clustering and frequency threshold / safety criteria, and further determine standard actions by learning sequence and dependency patterns and verifying the degree of agreement of adjacent steps.
[0052] For example, the video automatic generation device (200) can determine the camera viewpoint, playback speed, step decomposition, and emphasis method for each determined standard motion based on statistical values and preferences of the learning data, and can automatically generate animation clips in a physics engine-based 3D environment according to the determined parameters.
[0053] For example, the video automatic generation device (200) can identify step boundaries and key action candidates in the procedure description of a target manual provided from a user terminal (100), and generate a product description video by selecting and combining the optimal clips for each step according to a similarity index that combines semantic similarity, sequence suitability, and context suitability.
[0054] For example, the video automatic generation device (200) provides the generated product description video to the user terminal (100) and, if necessary, can optimize the naturalness of the sequence and playback time by applying transition patterns, timing control, and duplicate removal.
[0055] This image automatic generation device (200) may include a processor (201) and a memory (202).
[0056] The processor (201) performs major operations such as data collection and structuring, standard operation definition by pattern learning and verification, animation parameter determination, clip creation by 3D simulation and rendering, matching and combination, and result provision.
[0057] The processor (201) may include at least one of, for example, a microprocessor, a central processing unit (CPU), a GPU / NPU-based accelerator, a multicore / multiprocessor, an ASIC, or an FPGA.
[0058] The memory (202) can store a training dataset, a standard action dictionary, a parameter profile, generated animation clips and metadata, and program / model parameters executable by the processor (201).
[0059] The memory (202) may include, for example, volatile memory such as DRAM and non-volatile memory such as flash / SSD.
[0060] Such an automatic image generation device (200) may be implemented as a single physical server, but according to the embodiment, it may be configured to be distributed and processed in parallel across multiple computing nodes (e.g., GPU / NPU acceleration nodes, storage nodes, orchestration nodes) in a cloud environment, or implemented as a combination thereof.
[0061] For example, data collection and normalization processing can be deployed on nodes with high I / O scalability, standard operation definition and parameter determination on computation acceleration nodes, and clip creation, rendering, and sequence synthesis on graphics acceleration nodes, operating as microservices; and can be scaled to handle high volume requests through container-based autoscaling.
[0062] In this way, unlike conventional manual, template-centered production, the system (10) according to one embodiment of the present invention can automatically provide explanatory videos of accurate and consistent quality for various products and situations by combining data-based definition of standard operations with parameterized visualization and adaptive matching.
[0064] FIG. 2 is a flowchart illustrating a method for automatically generating product description videos using learning-based standard motion automatic extraction according to an embodiment of the present invention.
[0065] One embodiment illustrated in FIG. 2 is performed in an image automatic generation device (200).
[0066] In step S210, the image automatic generation device (200) can build a learning dataset based on multiple product manuals collected for multiple industry groups.
[0067] In one embodiment, the image automatic generation device (200) can collect product manuals from multiple industrial groups such as electronic products, furniture, automobiles, and medical devices, and convert the collected manuals into a structured standard format in procedural units to integrate the formats.
[0068] In one embodiment, the procedure unit standard format may include fields of [step identifier, action verb, target part / tool, caution phrase, illustration caption, reference image / frame path].
[0069] For example, manuals of different formats such as PDF, web pages, and video subtitles can be parsed and normalized into a step sequence like "1) bracket installation, 2) pipe connection, 3) electrical wiring, 4) test run".
[0070] In step S220, the image automatic generation device (200) can extract common patterns that are commonly repeated in multiple product manuals from a training dataset and define standard operations based on statistical verification of the common patterns.
[0071] In one embodiment, the image automatic generation device (200) can extract action verbs from a product manual and perform semantic clustering on the extracted action verbs to generate candidate patterns. Among the generated candidate patterns, a candidate pattern that satisfies a safety verification criterion while satisfying a preset frequency threshold can be defined as a standard action.
[0072] In one embodiment, the image automatic generation device (200) learns a pattern of sequential dependency between operations to calculate the dependency between the previous step and the next step for a candidate pattern, and if the dependency is greater than or equal to a preset threshold value, the candidate pattern can be defined as a standard operation.
[0073] For example, if the verb cluster "tighten / fasten / fix" satisfies the frequency threshold and safety criteria, and the degree of concordance of adjacent steps "align - fasten - check torque" is above the standard, it can be confirmed as a standard operation [fasten].
[0074] In step S230, the video automatic generation device (200) determines animation parameters for improving the understanding of the product description video for standard operation and can generate an animation clip based on the animation parameters.
[0075] Here, an animation clip refers to any visual object created and utilized to explain motion, and may include video clips, still / half-still frames, overlays such as text, icons, highlights, and guidelines, graphics synchronized with subtitles / narration, and composite product image results.
[0076] In one embodiment, animation parameters may include a camera viewpoint, playback speed, level of step breakdown, and an emphasis method. The video automatic generation device (200) may select a camera viewpoint for a standard motion from a set of candidate viewpoints based on the statistical derivation results of the training data. For example, the video automatic generation device (200) may set a playback speed based on the statistical values of the training data, determine a level of step breakdown based on the derived level of detail, and select an emphasis method based on the derived preference.
[0077] In one embodiment, the video automatic generation device (200) can automatically generate animation clips by applying determined animation parameters and simulating standard motions in a 3D environment with a physics engine applied. Visual parameters can be corrected to maintain consistency in visual style between the generated animation clips.
[0078] In one embodiment, the video automatic generation device (200) can generate animation clips in multiple versions and store each version in a clip library by assigning metadata of product type and parameter values.
[0079] For example, for the 'Pipe Connection' standard action, "45-degree oblique view - 1.2x speed - 3-step breakdown - guideline highlighting" is selected, allowing you to simultaneously create two version clips: one for beginners (zoomed in / slow speed / text reinforcement) and one for experts (points-oriented / abbreviated).
[0080] In step S240, the video automatic generation device (200) can adaptively calculate the correspondence between the target manual and the standard operation, and generate a product description video by combining animation clips based on the calculated correspondence.
[0081] Here, the term "target manual" refers to explanatory materials (including electronic documents, web pages, tables, and diagrams) describing the installation, assembly, initial setup, daily use, and maintenance procedures of the target product.
[0082] In one embodiment, the video automatic generation device (200) can identify step boundaries from the procedure description of the target manual and extract key action candidates for each step. For each step, a similarity index is calculated based on semantic similarity, order suitability, and context suitability between the key action candidate and the standard action, and the animation clip with the maximum similarity index can be selected as the optimal clip for that step.
[0083] In one embodiment, if the similarity index is below a threshold, replacement clips can be supplemented and selected in the upper and lower adjacent steps of the standard operation to reduce sequence gaps.
[0084] For example, the steps "Dock installation - Main unit charging - Map generation - Cleaning mode selection" are extracted from the robot vacuum cleaner initial setup manual, and a sequence can be configured by selecting the clip with the highest indicator for each step.
[0085] In this way, one embodiment illustrated in FIG. 2 can generate a product description video (S240) by adaptively calculating the correspondence between the target manual and the standard operation through the construction of a learning dataset (S210), the extraction of common patterns and statistical verification of the standard operation (S220), the determination of animation parameters and the generation of an animation clip (S230), and the automatic provision of description videos of consistent quality for various industry groups and product types.
[0087] FIG. 3 is a block diagram of a learning dataset construction pipeline within an image automatic generation device (200) according to one embodiment of the present invention.
[0088] Through the configuration illustrated in FIG. 3, the image automatic generation device (200) can build a learning dataset based on multiple product manuals collected for multiple industry groups. In the example illustrated in FIG. 3, the image automatic generation device (200) includes a collector (211), a parser (212), an extractor (213), a structurer (214), a standard format converter (215), and a dataset storage (216).
[0089] In this description, a standard format refers to a common schema for normalizing heterogeneous product manuals into procedural units. Furthermore, a procedural unit refers to the minimum set of sequential actions required to achieve a single objective within a product manual.
[0090] The collector (211) can collect product manuals in large quantities from multiple industrial sectors, such as electronic products, furniture, automobiles, and medical devices. Data sources may consist of manufacturer customer support web pages, public document repositories, corporate internal document servers, video subtitle files, etc. The collection method may combine batch collection and incremental collection, which updates only the changes. When collecting, the file format and source metadata are recorded together so that they can be used for quality control later.
[0091] The parser (212) can convert manuals of different formats, such as PDFs, web pages, scanned images, and video subtitles, into machineable text and layout structures. For example, the parser (212) can apply OCR to extract text containing characters and diagram captions included in an image. The parser (212) can preserve the hierarchical structure of tables and lists to improve the accuracy of subsequent procedure boundary recognition.
[0092] The extractor (213) can extract step candidates, action verbs, part / tool references, warning / prohibition phrases, and reference image location information from paragraphs, tables, and diagram captions. Here, an action verb refers to a verb or verb phrase indicating an action that the user must perform.
[0093] In one embodiment, the extractor (213) can detect the start and end of a step candidate using a list marker, a number, a conjunction, and a heading level. For example, each item in a numbered list such as "1. Bracket installation - 2. Pipe connection - 3. Electrical wiring - 4. Test run" can be extracted as a step candidate.
[0094] The structuring unit (214) can group the extraction results into procedural units and assign a step identifier, action verb, target part / tool, warning phrase, illustration caption, and reference image / frame path to each procedural unit. The structuring unit (214) can preserve sequence information and dependency relationships between steps so that they can be used for learning precedence dependency patterns in subsequent standard action definition steps.
[0095] A standard format converter (215) can convert structured procedure data into a standard format to integrate the format. For example, the field schema can be configured as [step identifier, action verb, target part - tool, warning phrase with assigned warning level, reference illustration - image - frame path, source metadata, language, version].
[0096] When a multilingual manual is input, the standard format converter (215) can ensure term consistency through term normalization and translation processing. For example, "fasten screw" in the English manual and "screw fastening" in the Korean manual can be mapped to the same action verb dictionary to be aggregated as the same pattern in subsequent learning.
[0098] The dataset storage (216) can store procedural data converted into a standard format. For example, the dataset storage (216) can be indexed by document - product - model - language - version keys when storing and can eliminate duplicates by assigning hash values at each stage.
[0099] In one embodiment, the dataset repository (216) can generate a quality report by detecting missing fields, abnormal order, and non-normal terms through a data validation pipeline.
[0100] In this way, the image automatic generation device (200) can execute the entire process of collection - parsing - extraction - structuring - standardization as a batch workflow and calculate statistical indicators for each stage (e.g., number of extracted action verbs, distribution of stage lengths, ratio of warning phrases, normal processing rate by language, etc.).
[0101] For example, a high proportion of action verbs related to fastening is observed in bookshelf assembly manuals within the furniture category, while a higher distribution of warning grades may be observed in medical devices compared to other categories. These differences can be utilized in subsequent standard operation definitions and the tuning of safety verification criteria.
[0102] According to the embodiment, the image automatic generation device (200) may be implemented in an on-premises single-server configuration or a cloud-based distributed configuration. When collecting in bulk, the collector (211) and the parser (212) are placed on an expansion node, and the subsequent steps after the extractor (213) can be processed in parallel on a node equipped with acceleration equipment. By adopting a metadata-centric storage structure, the device can be expanded without replacing the schema even when a new product family is added.
[0103] In this way, a learning dataset can be constructed through an embodiment illustrated in FIG. 3, and high-quality inputs necessary for common pattern extraction, standard motion definition, animation parameter determination, and clip generation can be consistently provided.
[0105] FIG. 4 is a flowchart illustrating a standard operation automatic definition procedure according to one embodiment of the present invention.
[0106] In step S410, the image automatic generation device (200) can extract action verbs from procedural text included in the training dataset.
[0107] In step S420, the image automatic generation device (200) can generate candidate patterns by performing semantic clustering on the extracted action verbs. Here, semantic clustering refers to a process of clustering action verbs that are semantically close based on synonym / synonym relationships or embedding similarity, and candidate patterns refer to patterns that are likely to be defined as standard actions as action bundles obtained as a result of clustering.
[0108] For example, "tighten," "sign," and "fix" can be grouped into the same cluster to generate candidate patterns.
[0109] In step S430, the image automatic generation device (200) can define a candidate pattern among the generated candidate patterns that satisfies a safety verification criterion while satisfying a preset frequency threshold as a standard operation.
[0110] Here, safety verification criteria refer to judgment criteria that include the coexistence of caution and warning statements in the manual, indicators of correlation with the prevention of safety-related accidents, and the results of mapping safety regulations by industry group. For example, if a fastening action is observed with sufficient frequency and consistently described along with safety caution statements, it can be defined as a standard action [fastening].
[0111] In one embodiment, the image automatic generation device (200) may define a case in which a frequency threshold and a safety verification criterion are simultaneously satisfied among candidate patterns as a standard operation.
[0112] In one embodiment, when calculating a frequency threshold, the image automatic generation device (200) may determine that the occurrence rate of candidate patterns for each comparison group satisfies a predetermined minimum rate and simultaneously statistically significantly exceeds the average occurrence rate of the group as a target for standard operation.
[0113] Specifically, the image automatic generation device (200) can determine the group to which the candidate pattern belongs and select documents of that group. For example, an industry group and a language can be set as one group, such as Air Conditioner Installation - Korean.
[0114] Afterwards, the image automatic generation device (200) calculates the ratio of the total documents to the documents in which a candidate pattern appears within the group, and based on this, can calculate the average and standard deviation of the occurrence rates for all candidate patterns in the group.
[0115] Subsequently, the image automatic generation device (200) can determine whether the frequency threshold is satisfied by applying the following dual conditions. First, the occurrence rate of candidate patterns may exceed the minimum standard rate per group. Second, it may be determined that the occurrence rate of candidate patterns significantly exceeds the group average. Here, the second condition can be implemented using a standard score standard based on the standard deviation or a conservative lower limit standard.
[0116] In one embodiment, the minimum standard ratio for the general household appliance group may be set to 20%, and for the medical device group to 30%, and the prominence standard relative to the average may be set to a standard score of 1.0 or higher.
[0117] For example, if a transaction pattern is observed in about 60% of the documents in the group, and the average of the group is about 30% and satisfies the standard score criteria, the transaction pattern can be determined as a target that meets the frequency threshold.
[0118] In one embodiment, when calculating safety verification criteria, the image automatic generation device (200) evaluates whether a candidate pattern has sufficient safety context using multiple signals and can determine that a case satisfying a group-specific criterion value is a target for standard operation.
[0119] Specifically, the image automatic generation device (200) can determine the group to which the candidate pattern belongs and select the manual text and table - picture caption - icon - annotation of the group as targets for extraction of safety-related clues.
[0120] For example, the image automatic generation device (200) can collect the following safety signals and normalize them between 0 and 1.
[0121] - Warning Sign Coexistence Ratio: The ratio of warning signs or safety icons, such as Caution, Warning, and Danger, listed together with the candidate pattern at the same level can be calculated.
[0122] - Warning Grade Weight: The grades assigned in the document, such as Minor, Caution, Warning, and Danger, can be scored to calculate an average value.
[0123] - Antecedent-Post-Dependency Stability: The degree to which candidate patterns are repeatedly combined in a consistent order with preceding and succeeding steps can be calculated using step sequence statistics. Here, antecedent-post-dependency stability refers to the lower bound statistic of the degree of agreement in which a specific step appears together with its superior and inferior adjacent steps within the same group of manuals.
[0124] - Protection / Blocking Keyword Coexistence Ratio: The ratio of simultaneous appearance of keywords related to protective equipment and blocking procedures, such as gloves, safety glasses, power cutoff, and valve closure, can be calculated.
[0125] In one embodiment, the image automatic generation device (200) may set weights for the above signals according to an industry-group policy table and calculate a safety score by multiplying the normalized values by the weights and summing them. The weights may be adjusted according to the domain risk level. For example, the electrical and medical device group may have the weights for warning grades and protection / blocking items increased, and the furniture assembly group may have the weights for priority dependency stability increased relatively.
[0126] In one embodiment, the image automatic generation device (200) may set a reference value for each group and determine that the safety verification criteria are satisfied when the safety score is above the reference value. The reference value may be set as a lower quantile of the data distribution or a fixed value of the operation policy. If the sample size is small or the document quality is low, the score may be conservatively adjusted by reflecting the reliability of each signal. For example, regarding a candidate pattern called "fastening," if warning markers are accompanied by a high proportion, the sequential combination of alignment, fastening, and torque verification is stably repeated, and there is sufficient mention of gloves and safety glasses, a safety score exceeding the reference value of the home appliance group, which has a lower risk than electrical or medical devices, may be calculated. In this case, the fastening candidate pattern may be determined to satisfy the safety verification criteria. Conversely, a candidate pattern with almost no warning markers, insufficient mention of protection and blocking procedures, and unstable step combination may be withheld for failing to meet the reference value.
[0127] In step S440, the image automatic generation device (200) can learn the pattern of precedence dependency between operations and calculate the dependency between the previous step and the next step for the candidate pattern.
[0128] A precedence dependency pattern refers to a statistical relationship that reflects the sequence of procedural steps and the necessary conditions between steps; here, dependency is a value that quantifies the tendency of a specific candidate pattern to appear together with preceding and succeeding steps, representing an indicator of adjacent step agreement.
[0129] In one embodiment, the image automatic generation device (200) can normalize the procedure sequence into candidate pattern units and aggregate consecutive step pairs per document to calculate a dependency index.
[0130] For example, an image automatic generation device (200) can calculate a dependency index by combining four elements: transition support, order consistency, proximity weighting, and contextual consistency. Transition support may refer to the ratio of a specific transition (from u to v) observed in different documents, and order consistency refers to the ratio of u - v order being maintained without being reversed. Proximity weighting can be implemented as a function that assigns greater weight as the distance between steps is closer, and contextual consistency refers to the degree to which common part names, tool names, cautionary terms, and warning signs are listed together in two steps. The four elements can be combined as a weighted sum, and a conservative lower bound can be applied to groups with a small sample size to suppress overestimation.
[0131] If the dependency indicator is above the threshold value for each group, the corresponding transition can be determined as having an established dependency. For example, if alignment followed by fastening is observed at a high rate in multiple manuals, cases of reversal are rare, the two steps mainly appear as adjacent cells, and common part names and cautionary phrases are listed consecutively, the alignment-fastening transition can be calculated as having a dependency that stably exceeds the threshold value. Conversely, transitions like alignment-lubrication, where the order frequently changes across documents, the distance between steps is large, and contextual consistency is low, may be determined to be below the threshold.
[0132] In one embodiment, the image automatic generation device (200) can adjust reference values and weights by considering document length and risk level by industry group. For example, the medical device group can have the weights for order consistency and context consistency increased, and the furniture assembly group can have the influence of proximity weight increased to more strictly reflect the actual workflow.
[0133] In step S450, the image automatic generation device (200) may define the corresponding candidate pattern as a standard operation if the dependency calculated in step S440 is greater than or equal to a preset threshold value. The threshold value can be set based on data according to the sample size and variability of each industry group.
[0134] Subsequently, the image automatic generation device (200) can register defined standard actions in a standard action dictionary and assign metadata to each standard action so that it can be referenced in subsequent steps. The metadata may include a set of representative action verbs, example phrases, safety ratings, a summary of precedence dependencies, applicable industries, example parts, and tool information.
[0135] In one embodiment, the image automatic generation device (200) can reduce excessive generalization or omission by applying frequency thresholds and safety verification standards differently according to industry group. For example, in a medical device manual, the reliability of warning phrases can be increased, and in a furniture assembly manual, the frequency threshold of fastening-related operations can be raised to increase the stability of standard operations.
[0136] In one embodiment, the image automatic generation device (200) may selectively apply a sequence model or a graph-based frequency accumulation method to the learning of precedence dependencies. In the semantic clustering step of candidate patterns, domain-specific expressions may be incorporated by combining a pre-constructed action verb dictionary and an embedding-based proximity search.
[0137] In this way, through an embodiment illustrated in FIG. 4, reliably standardizing common repetitive actions in manuals of various industries is possible, and consistent quality can be ensured in determining animation parameters and matching target manuals in subsequent steps.
[0139] FIG. 5 is a flowchart illustrating a procedure for determining animation parameters according to an embodiment of the present invention.
[0140] In step S510, the image automatic generation device (200) can receive a standard operation defined in FIG. 4 as input and load an initial setting for determining animation parameters.
[0141] As in the example above, animation parameters refer to configuration values including the camera viewpoint, playback speed, step breakdown level, and emphasis method used to visually represent standard motion.
[0142] In step S520, the image automatic generation device (200) can select a camera viewpoint for a standard operation from a set of candidate viewpoints based on the statistical derivation result of the training data.
[0143] The set of candidate viewpoints can be predefined as front, oblique top, magnified adjacent, sectional perspective, etc. For example, for a standard fastening operation, an oblique top close-up viewpoint can be selected to clearly show the contact between the screw head and the tool.
[0144] In step S530, the video automatic generation device (200) can set the playback speed of the standard operation based on the statistical value of the training data.
[0145] For example, the video automatic generation device (200) can distribute the basic speed and deceleration sections by considering the distribution of user proficiency, the average viewing break-off section, and whether safety warning messages are included. For example, the danger section can be decelerated by 0.75 times, and the repetition section can be accelerated by 1.25 times.
[0146] In step S540, the image automatic generation device (200) can determine the level of step decomposition of the standard operation based on the level of refinement derived from the training data.
[0147] Here, the level of subdivision refers to the tendency to describe the same standard operation by generally dividing it into several sub-steps, and the level of step decomposition refers to the extent to which a single standard operation is divided into several detailed frames for explanation. For example, a fastening operation can be decomposed into three steps: alignment preparation, initial fastening, and torque verification.
[0148] To this end, the image automatic generation device (200) can derive a recommended level of subdivision for each standard operation by statistically aggregating the list depth, subheading unit, command boundary, warning phrase insertion location, illustration caption transition point, etc., in the procedural structure of the training data, and based on this, the current standard operation can be broken down into at least one of 2 stages, 3 stages, or 4 stages.
[0149] For example, if it is observed in the training data that the subheadings and warning messages for the standard fastening operation are mainly repeated at three points: alignment preparation, initial fastening, and torque verification, the step decomposition level can be determined as 3.
[0150] Conversely, in product families where assembly jigs are provided, if a tendency is observed for the list depth to be shallow and the statements to be described consecutively, the same fastening operation can be reduced to two steps of alignment and fastening.
[0151] In addition, if the insertion of supplementary explanations or enlarged diagrams, which can be seen as signals of reduced understanding, appears frequently immediately after a specific sub-step in the training data, the level of step decomposition can be increased by adding that point as a decomposition boundary.
[0152] In step S550, the image automatic generation device (200) can select a method of emphasizing standard actions based on preferences derived from training data.
[0153] In one embodiment, the emphasis method may include at least one of border highlighting, translucent masking, guideline, flashing icon, and text caption. For example, a warning color border and a caution icon may be added in a danger section, and a guideline may be activated in a complex section.
[0154] Afterwards, the video automatic generation device (200) can automatically generate an animation clip by applying determined animation parameters and simulating standard movements in a 3D environment with a physics engine applied. This will be described later with reference to FIG. 6.
[0155] In this way, an embodiment illustrated in FIG. 5 can receive standard motions as input and determine animation parameters such as camera viewpoint, playback speed, step breakdown level, and emphasis method based on data, and the determined animation parameters can be used to generate animation clips for product description video synthesis in the subsequent procedure of FIG. 6.
[0157] FIG. 6 is a flowchart illustrating the animation clip creation and library management procedure according to one embodiment of the present invention.
[0158] The animation parameters determined in Fig. 5 can be applied at each step of Fig. 6.
[0159] In step S610, the image automatic generation device (200) can load standard operations and animation parameters determined in FIG. 5, and prepare 3D assets and physics engine settings required for simulation.
[0160] For example, 3D assets refer to target product models, part-tool models, material-texture information, lighting rigs, camera rigs, etc., and physics engine settings refer to gravity, friction coefficients, moments of inertia, collision handling, joint constraints, etc.
[0161] In step S620, the automatic image generation device (200) can generate a basic animation sequence by simulating standard movements in an environment where a physics engine is applied. Here, the camera viewpoint, playback speed, level of step resolution, and emphasis method may be determined from FIG. 5. For example, in the case of a standard fastening movement, the contact between the screw and the tool, rotation, and torque increase sections can be reproduced with dynamics similar to reality.
[0162] In step S630, the video automatic generation device (200) can composite a visual overlay onto the base sequence. The overlay may include at least one of a highlight border, a translucent mask, a guideline, a blinking icon, and a text caption. It can be composited to add a warning color border and a caution icon in dangerous sections, and to activate the guideline in complex sections.
[0163] In step S640, the video automatic generation device (200) can correct the visual style of the generated animation clip to match the project standards. The correction targets may include color palette, lighting intensity, shadow density, texture resolution, subtitle font size, margin rules, etc. Automatic correction can be performed so that the sense of incongruity is minimized even when clips generated from different standard movements are combined into a single video.
[0164] In step S650, the video automatic generation device (200) can generate multiple versions for the same standard operation. The multiple versions can be distinguished by a combination of user proficiency level, viewing angle, playback length, whether text readability is enhanced, and accessibility options. For example, a version for beginners can be generated with zoom view - low speed - high step breakdown - enhanced captioning, and a version for experts can be generated with point view - basic speed - medium step breakdown - minimal captioning.
[0165] In step S660, the image automatic generation device (200) may assign metadata for each generated version. For example, the metadata may include a standard motion ID, product family - model, camera viewpoint, playback speed, step breakdown level, highlighting method, time of generation, engine version, language, and accessibility tag.
[0166] Afterward, the video automatic generation device (200) can store animation clips and metadata in a clip library. When storing, indexing with a project - product - standard action - version key and performing hash-based duplicate detection can prevent duplicate storage of the same asset. By generating a preview thumbnail and a low-resolution proxy together, the responsiveness of the subsequent search - matching step can be improved.
[0167] According to an embodiment, the image automatic generation device (200) may perform further quality inspections. The inspection may include screen readability, color vision affinity, subtitle-motion synchronization, frame missing-flickering, maximum length compliance, file integrity, etc. Items that do not meet the criteria may be automatically readjusted or given a hold flag.
[0168] According to an embodiment, the image automatic generation device (200) may automatically perform regeneration and re-correction of failure items according to an operation policy, or transfer them to a review queue. For example, if the synchronization error is below a threshold, automatic correction may be performed, and if it is above a threshold, it may be branched to a review queue.
[0169] In this way, one embodiment illustrated in FIG. 6 can systematically generate and manage animation clips of standard motion units by applying parameters determined in FIG. 5, and can be immediately utilized in subsequent target manual matching and product description video combination steps.
[0171] FIG. 7 is a flowchart illustrating an adaptive matching and animation clip selection procedure between a target manual and a standard operation according to an embodiment of the present invention.
[0172] In step S710, the image automatic generation device (200) can receive the procedure description of the target manual as input, identify the step boundaries, and extract key action candidates for each step. As previously mentioned, the target manual refers to explanatory material describing the installation, assembly, initial setup, daily use, and maintenance procedures of the target product. Key action candidates refer to a set of action expressions derived from a verb or verb phrase that directs an action within a step sentence.
[0173] In one embodiment, the image automatic generation device (200) can detect step boundaries by featuring a number list, headings, conjunctions, graphic caption clues, and whether there is an imperative form within the sentence.
[0174] In step S720, the image automatic generation device (200) can calculate semantic similarity, order suitability, and context suitability between a core action candidate and a standard action for each step. Semantic similarity refers to a value calculated in the embedding space regarding the semantic proximity between a core action candidate and a standard action representation, and order suitability refers to the degree to which the context of the current step matches a typical adjacent step of the standard action by referring to the precedence dependency index calculated in FIG. 4. Additionally, context suitability refers to the degree to which the part-tool-caution-icon information included in the step text matches the standard action metadata.
[0175] In one embodiment, the image automatic generation device (200) can normalize three indicators and calculate a similarity indicator as a weighted sum.
[0176] In step S730, the video automatic generation device (200) can select the animation clip with the maximum similarity index as the optimal clip for that step.
[0177] In one embodiment, the video automatic generation device (200) may provisionally select the top N candidate clips and then reorder the final order by applying a penalty so that the combination of length-time-highlight method does not change abruptly between consecutive steps. For steps below the threshold, alternative clips may be supplemented and selected from the upper-lower adjacent steps of the standard operation to reduce sequence gaps.
[0178] In step S740, the video automatic generation device (200) can verify the naturalness and readability of the transitions of the selected animation clips and, if necessary, perform fine-tuning of the camera viewpoint, playback speed, and emphasis method. The transition verification can be automatically evaluated based on the amount of change in scene brightness, the amount of movement of the subject center, subtitle-motion synchronization error, etc.
[0179] In one embodiment, the image automatic generation device (200) can be adjusted to reduce the amount of viewpoint movement when the same part appears repeatedly in consecutive steps, and to place a warning icon and a highlight in advance at the start frame of the danger section.
[0180] In step S750, the video automatic generation device (200) can determine the optimal clip and correction parameters for each step as a combination plan and transmit it to a subsequent compositing step along with timeline specifications. Here, the combination plan refers to a set of selected clip IDs, start-end timecodes, transition methods, and subtitle-overlay placement information.
[0181] In one embodiment, the image automatic generation device (200) may include an alternative palette - enlarged caption - audio description track use flag in the combination plan when accessibility requirements are enabled.
[0182] For example, if the steps of dock installation, main body charging, map generation, and cleaning mode selection are extracted from the initial setup manual of a robot vacuum cleaner, the automatic image generation device (200) can prioritize selecting clips with high semantic similarity for each step, and in sections requiring prior dependencies such as alignment, connection, and torque verification, clips with high sequence suitability can be selected by adding points. If a warning message regarding power connection is emphasized during the main body charging step, clips with high context suitability can be finally selected.
[0183] In this way, one embodiment illustrated in FIG. 7 can identify step boundaries in a target manual and extract key action candidates, then select an optimal clip based on a similarity index of semantic similarity - order suitability - context suitability, and finalize a combination plan by correcting transition quality and readability, which can be reliably utilized in a subsequent product description video synthesis step.
[0185] FIG. 8 is a flowchart illustrating the procedure for providing final synthesis and rendering of a product description video according to an embodiment of the present invention.
[0186] The combination plan finalized in Fig. 7 can be referenced in each step of Fig. 8.
[0187] In step S810, the video automatic generation device (200) can load a combination plan and place the selected animation clip on the timeline.
[0188] In one embodiment, the combination plan may include a set of clip IDs selected stepwise, start and end timecodes, transition methods, and subtitle overlay placement information.
[0189] The video automatic generation device (200) can check the gaps between clips and, if necessary, fine-tune the transition length and start frame to align the timeline.
[0190] In step S820, the video automatic generation device (200) can apply transitions, subtitles, and overlays to the timeline. Transitions may be limited to presets according to project rules, such as cuts, dissolves, and wipes. For example, subtitles and overlays are aligned based on the placement information specified in FIG. 7, and can be adjusted to minimize the amount of viewpoint movement and the change in overlay position in sections where the same part appears consecutively.
[0191] In step S830, the video automatic generation device (200) can synchronize narration, sound effects, and background sound sources. The narration can be synthesized based on text or matched with a prepared sound source. In one embodiment, the video automatic generation device (200) can align narration keyframes to the step start event and superimpose a short tone of warning sound in the danger section to draw attention. The subtitle-narration synchronization error can be automatically corrected within an acceptable range.
[0192] In step S840, the video automatic generation device (200) can configure alternative tracks that reflect accessibility and multilingual requirements. Accessibility modes may include a color vision deficiency friendly palette, enlarged captions, and audio description tracks. Multilingual subtitles may place text translated into the target language according to the timecode.
[0193] In step S850, the image automatic generation device (200) can perform rendering according to a predefined encoding profile. The encoding profile may include resolution, frame rate, bit rate, and codec.
[0194] In one embodiment, a proxy for internal review (e.g., 720p) and a master for distribution (e.g., 1080p or 4K) are generated simultaneously, and the color space and metadata are embedded together to prevent color distortion during re-encoding.
[0195] In one embodiment, the video automatic generation device (200) may package and provide outputs and metadata. The metadata may include a target product ID, version, a list of standard actions used, accessibility support items, language, length, and time of creation. For example, the video automatic generation device (200) may provide a preview link to a user terminal and, if necessary, upload it to a content management system repository.
[0196] In one embodiment, the video automatic generation device (200) may generate a quality inspection log. Inspection items may include subtitle-narration synchronization error, frame drop rate, whether contrast-readability standards are met, transition abruptness indicators, etc. If an item that does not meet the standards is detected, the corresponding section of the combination plan may be marked and recorded so that re-synthesis or correction is possible.
[0197] For example, when the dock installation, main body charging, map generation, and cleaning mode selection clips selected in FIG. 7 are placed on the timeline for the robot vacuum cleaner initial setup sequence, the video automatic generation device (200) can arrange a warning icon and narration at the start frame of the main body charging step and set the transition to the map generation step as a dissolve to compose the scene change smoothly. When accessibility mode is enabled, an alternative version including enlarged captions and audio description tracks can be output in parallel.
[0199] The device described above may be implemented as a hardware component, a software component, and / or a combination of a hardware component and a software component. For example, the device and components described in the embodiments may be implemented using one or more general-purpose or special-purpose computers, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing unit may execute an operating system (OS) and one or more software applications executed on said operating system. Additionally, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing unit may be described as being used as a single unit, but those skilled in the art will understand that the processing unit may include multiple processing elements and / or multiple types of processing elements. For example, the processing unit may include multiple processors or one processor and one controller. In addition, other processing configurations, such as parallel processors, are also possible.
[0200] Software may include computer programs, code, instructions, or a combination of one or more of these, and may configure a processing unit to operate as desired or command the processing unit independently or collectively. Software and / or data may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave so as to be interpreted by the processing unit or to provide instructions or data to the processing unit. Software may be distributed over networked computer systems and may be stored or executed in a distributed manner. Software and data may be stored on one or more computer-readable recording media.
[0201] The method according to the embodiment may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program instructions, data files, data structures, etc., either alone or in combination. The program instructions recorded on the medium may be those specifically designed and configured for the embodiment, or they may be those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc. The hardware devices described above may be configured to operate as one or more software modules to perform the operation of the embodiment, and vice versa.
[0202] Although the embodiments have been described above with reference to limited examples and drawings, those skilled in the art can make various modifications and variations from the description above. For example, suitable results can be achieved even if the described techniques are performed in a different order than described, and / or the components of the described system, structure, device, circuit, etc. are combined or assembled in a form different from described, or replaced or substituted by other components or equivalents.
[0203] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims set forth below.
[0204] Although specific embodiments have been described in the detailed description of this document, it will be obvious to those skilled in the art that various modifications are possible within the scope of this document. Explanation of the symbols
[0205] 10: Automatic Product Description Video Generation System 100 : User terminal 200 : Automatic image generation device
Claims
Claim 1 A method for automatically generating a product description video using learning-based standard action automatic extraction, comprising: a step of constructing a learning dataset based on multiple product manuals collected for multiple industry groups; a step of extracting a common pattern that is repeated in common across the multiple product manuals from the learning dataset and defining a standard action based on statistical verification of the common pattern; a step of determining animation parameters for the standard action and generating an animation clip based on the animation parameters; and a step of adaptively calculating a correspondence relationship between a target manual and a standard action and generating a product description video by combining animation clips based on the calculated correspondence relationship; wherein the step of defining the standard action includes: a step of extracting action verbs from the product manuals and generating candidate patterns by performing semantic clustering on the extracted action verbs; and a step of defining a candidate pattern among the generated candidate patterns that satisfies a safety verification criterion while satisfying a preset frequency threshold as the standard action. Claim 2 A method for automatically generating product description videos using learning-based standard operation automatic extraction, wherein the step of constructing the learning dataset comprises: a step of collecting the plurality of product manuals from the plurality of industry groups; and a step of converting the plurality of product manuals into a structured standard format in procedural units to perform format integration. Claim 3 delete Claim 4 A method for automatically generating product description videos using learning-based standard action automatic extraction, wherein the step of defining the standard action further comprises: a step of learning a pattern of sequential dependencies between actions in a product manual to calculate the dependency between the immediate and subsequent steps for the candidate pattern; and a step of defining the candidate pattern as a standard action if the dependency is greater than or equal to a preset threshold value. Claim 5 A method for automatically generating a product description video using learning-based standard motion automatic extraction, comprising: a step of constructing a learning dataset based on multiple product manuals collected for multiple industry groups; a step of extracting a common pattern that is commonly repeated in the multiple product manuals from the learning dataset and defining a standard motion based on statistical verification of the common pattern; a step of determining animation parameters for the standard motion and generating an animation clip based on the animation parameters; and a step of adaptively calculating a correspondence relationship between a target manual and a standard motion and combining animation clips based on the calculated correspondence relationship to generate a product description video; wherein the animation parameters include a camera viewpoint, playback speed, step decomposition level, and emphasis method; and the step of determining the animation parameters includes: a step of selecting a camera viewpoint for the standard motion from a set of candidate viewpoints based on the statistical derivation result of the learning data; a step of setting the playback speed of the standard motion based on the statistical value of the learning data; a step of determining the step decomposition level of the standard motion based on the level of granularity derived from the learning data; and a step of selecting an emphasis method of the standard motion based on the preference derived from the learning data. Claim 6 A method for automatically generating product description videos using learning-based standard motion automatic extraction, wherein the step of generating the animation clip further comprises: a step of automatically generating the animation clip by simulating the standard motion in a 3D environment with a physics engine applied by applying the determined animation parameters; and a step of correcting visual parameters to maintain consistency in visual style between the generated animation clips. Claim 7 A method for automatically generating product description videos using learning-based standard motion automatic extraction, wherein the step of generating the animation clip further comprises: generating the animation clip in a plurality of versions; assigning product type and parameter values as metadata to each version of the animation clip; and storing the plurality of versions of the animation clip and the metadata in a clip library. Claim 8 A method for automatically generating a product description video using learning-based standard action automatic extraction, comprising: a step of constructing a learning dataset based on multiple product manuals collected for multiple industry groups; a step of extracting a common pattern that is repeated in common across the multiple product manuals from the learning dataset and defining a standard action based on statistical verification of the common pattern; a step of determining animation parameters for the standard action and generating an animation clip based on the animation parameters; and a step of adaptively calculating a correspondence relationship between a target manual and a standard action and generating a product description video by combining animation clips based on the calculated correspondence relationship; wherein the step of generating a product description video by combining animation clips comprises: a step of identifying step boundaries from the procedural description of the target manual and extracting key action candidates for each step; a step of calculating a similarity index for each step based on semantic similarity, order suitability, and context suitability between the key action candidates and the standard action; and a step of selecting the animation clip with the maximum similarity index as the optimal clip for the corresponding step. Claim 9 A system for automatically generating product description videos using learning-based standard action automatic extraction, comprising: a user terminal providing a target manual; and an automatic video generation server that receives the target manual from the user terminal and automatically generates a product description video based on the target manual and provides it to the user terminal; wherein the automatic video generation server constructs a learning dataset based on multiple product manuals collected for multiple industry groups, extracts common patterns that are commonly repeated in the multiple product manuals from the learning dataset and defines a standard action based on statistical verification thereof, determines animation parameters for the standard action and generates an animation clip based on the animation parameters, adaptively calculates the correspondence relationship between the target manual and the standard action, and generates the product description video by combining the animation clips based on the correspondence relationship; and wherein, to define the standard action, the automatic video generation server extracts action verbs from the product manual, performs semantic clustering on the extracted action verbs to generate candidate patterns, and among the generated candidate patterns, defines a candidate pattern that satisfies safety verification criteria while satisfying a preset frequency threshold as the standard action. Claim 10 A computer-readable recording medium combined with hardware, storing a computer program for performing any one of the methods of paragraphs 1 to 2 and 4 to 8.
Citation Information
Patent Citations
Home appliances user guide video contents and advertisement providing method and system
KR1020170133110A
System for management and control of an enterprise
US20090327023A1