System for multi-layer verification and automatic correction of generative ai content using back-calculation of physical simulation data

KR103002930B1Active Publication Date: 2026-08-12CUBEBERRY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2026-01-27
Publication Date
2026-08-12

Smart Images

  • Figure 112026011043405-PAT00004_ABST
    Figure 112026011043405-PAT00004_ABST
Patent Text Reader

Abstract

A method for controlling content generation using a generative artificial intelligence model by a processor of a device according to some embodiments of the present disclosure may include: a step of generating a first data set by processing the input data when input data for content generation is received from a user; a step of determining whether the first data set satisfies the predefined generation conditions by comparing the first data set and the content to be generated in a vector space with predefined generation conditions; a step of generating a second data set including information on at least one of the location of an object or the operation of an object included in the content to be generated based on the first data set when it is determined that the first data set satisfies the predefined generation conditions; a step of determining whether the location of the object and the operation of the object satisfy predefined constraints; a step of extracting analysis data when it is determined that the location of the object and the operation of the object do not satisfy the predefined constraints; a step of converting the analysis data into control information in a form interpretable by the generative artificial intelligence model; and a step of generating final content by providing the control information as a conditional input to the generative artificial intelligence model.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present disclosure relates to a multi-layer verification and automatic correction system for generative AI content, and specifically to a method for multi-layered verification and automatic correction of the physical consistency of generative AI content by inversely calculating analysis data extracted from physical simulation results and converting it into control information for a generative AI model. Background Technology

[0002] With the recent advancement of artificial intelligence technology, generative AI, which automatically generates various forms of content such as text, images, and videos, is being widely utilized across the content industry. In particular, generative AI based on Large-Scale Language Models (LLM) and Diffusion Models is attracting attention for its ability to produce high-quality visual and narrative content with relatively simple input; however, it still harbors structural limitations in terms of the reliability and consistency of the generated results.

[0003] Conventional generative AI technology is based on a method of generating content by probabilistically predicting the next token or pixel according to input prompts. This leads to a problem where so-called physical hallucinations, which violate the physical laws of the real world, frequently occur during the generation process. For example, results that do not conform to physical causality are generated, such as parts of the human body penetrating objects or movements defying the laws of gravity or collision; this significantly degrades the realism and immersion of the content.

[0004] These physical illusions frequently manifest as a phenomenon where two objects penetrate each other, particularly in collision situations; the degree of penetration can be quantified as the penetration depth between the colliding objects. Although penetration depth is an indicator capable of numerically expressing the severity of physical law violations, conventional generative AI technologies have limitations in that they cannot recognize this physical penetration information during the generation process or control the generated results based on it.

[0005] Existing response methods to this mainly consist of manually correcting the erroneous results after a human check, or repeatedly regenerating the same prompt by changing the random seed. However, since these methods rely on simple repetition without analyzing the cause of the error or performing systematic correction, they involve significant waste of computational resources and carry the problem of rapidly increasing content production time and costs.

[0006] Furthermore, conventional technology suffers from the problem of failing to reliably maintain consistency in narrative elements such as character appearance, personality, behavioral patterns, and world-building. Contextual breakdown occurs, such as the same character appearing differently in each scene or exhibiting behavior unrelated to their established personality; this acts as a critical constraint in the production of world-building-based content or content requiring a continuous narrative structure.

[0007] Furthermore, large-scale language model-based story generation systems have limitations in continuously remembering and reflecting long-term narrative developments or complex worldview information. This leads to problems of narrative inconsistency, where connectivity between scenes deteriorates or characters' memories, goals, and emotional states are disconnected in the next scene. In particular, as the narrative length increases, initial settings or major events are forgotten, resulting in a disorganized plot and distorted causality.

[0008] To address this problem, some conventional technologies propose methods such as referencing external knowledge databases or retrospectively modifying and correcting generated results; however, these approaches are merely structures that discard or modify content containing errors after it has already been created. Consequently, significant computational resources are wasted when generating high-resolution video or multimodal content, and these methods have limitations in that they fail to fundamentally prevent the occurrence of errors.

[0009] Furthermore, conventional generative AI technologies fail to provide a systematic generative structure that comprehensively verifies and controls functions such as semantic consistency, physical consistency, and the maintenance of narrative and character consistency. Consequently, even if partial improvements are made to individual elements, the current situation remains insufficient to guarantee spatiotemporal consistency and reliability at the overall content level.

[0010] Therefore, there is an urgent need to develop a new content generation control method that can fundamentally prevent physical errors and the collapse of narrative consistency by comprehensively verifying the consistency of semantic settings, narrative context, and physical laws prior to the stage where generative AI generates content, and by generating content based only on data that has passed this verification. Prior art literature

[0011] Chinese Patent Application No. 2025-10020616 (Filed on January 7, 2025) The problem to be solved

[0012] The present disclosure aims to solve the aforementioned problems and other problems. The technical problem to be achieved by some embodiments of the present disclosure is to detect errors that violate physical constraints in advance or during the content generation process using a generative artificial intelligence model, and to automatically control and correct the generation process of said generative artificial intelligence model by inversely calculating analysis data extracted from physical simulation results and converting it into control information interpretable by said generative artificial intelligence model.

[0013] The technical problems to be solved in this disclosure are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art from the description below. means of solving the problem

[0014] A method for controlling content generation using a generative artificial intelligence model by a processor of an apparatus according to some embodiments of the present disclosure comprises: receiving input data for content generation from a user, processing said input data to generate a first data set including at least one of semantic information, structural characteristic information, or narrative context information regarding content to be generated; determining whether the first data set satisfies said predefined generation conditions by comparing said first data set and said content to be generated in a vector space; if it is determined that the first data set satisfies said predefined generation conditions, generating a second data set including information on at least one of the location of an object included in said content to be generated or the action of said object based on said first data set; determining whether the location of said object and the action of said object satisfy said predefined constraints by performing an evaluation according to said constraints on said second data set; if it is determined that the location of said object and the action of said object do not satisfy said predefined constraints, extracting analysis data including the degree of discrepancy with said predefined constraints and the location where the discrepancy occurred. The method may include the step of converting the analysis data into control information in a form interpretable by the generative artificial intelligence model; and the step of generating final content by providing the control information as a conditional input to the generative artificial intelligence model.

[0015] According to some embodiments of the present disclosure, the control information may include at least one of map data indicating the degree of discrepancy with the predefined constraint and the location where the discrepancy occurred, mask data generated based on the map data and specifying a part of the content to be generated, and weight data calculated based on the degree of discrepancy with the predefined constraint.

[0016] According to some embodiments of the present disclosure, the weight data is applied to a negative prompt tag or prompt element included in the input of the generative artificial intelligence model to control the degree to which the negative prompt tag or the prompt element is selected or suppressed during the generation process of the content to be generated.

[0017] According to some embodiments of the present disclosure, the step of determining whether the predefined generation condition is satisfied may include: a step of vectorizing the first data set and the predefined generation condition to calculate a similarity in the vector space; and a step of determining whether the predefined generation condition is satisfied based on whether the similarity satisfies a reference value.

[0018] According to some embodiments of the present disclosure, the step of determining whether the location of the object and the operation of the object satisfy predefined constraints may include: placing a simplified object that approximates the location of the object or the operation of the object in a virtual space; and determining whether the location of the object or the operation of the object satisfies the predefined constraints by performing an operation on the simplified object according to predefined physical constraints.

[0019] According to some embodiments of the present disclosure, the prior-defined physical constraint may include a condition for evaluating at least one of the possibility of collision between the simplified object and another object, a change in the position of the simplified object due to gravity, a restriction on the relative movement of the simplified object due to friction, or the continuity of motion of the simplified object due to inertia.

[0020] According to some embodiments of the present disclosure, the step of generating final content by providing the control information as a conditional input to the generative artificial intelligence model may include: generating guide data by extracting the spatial location of the object, the relative arrangement of the object, and the operational state of the object from the second data set; and generating the final content based on the input data by providing the control information as a conditional input to the generative artificial intelligence model together with the guide data and a previously registered reference image or appearance characteristic data.

[0021] According to some embodiments of the present disclosure, the guide data is generated corresponding to a plurality of time intervals, and the control information may be generated for each of the plurality of time intervals or may be generated in a form that is commonly applied to the plurality of time intervals.

[0022] According to some embodiments of the present disclosure, if it is determined that the first data set does not satisfy the predefined generation conditions, the method may further include a step of controlling so that subsequent steps for generating the final content are not performed.

[0023] According to some embodiments of the present disclosure, when it is determined that the location of the object and the operation of the object satisfy the predefined constraints, the method may further include the step of generating guide data by extracting the spatial location of the object, the relative placement of the object, and the operation state of the object from the second data set; and the step of generating the final content based on the input data by providing the guide data and a pre-registered reference image or appearance characteristic data as conditional inputs to the generative artificial intelligence model.

[0024] The technical solutions obtainable in this disclosure are not limited to the solutions mentioned above, and other solutions not mentioned will be clearly understood by those skilled in the art to which this disclosure belongs from the description below. Effects of the invention

[0025] The method for automatically controlling and correcting the generation process of a generative artificial intelligence model according to the present disclosure is described as follows.

[0026] According to some embodiments of the present disclosure, the generation process of generative AI content may be performed in a manner controlled stepwise after prior verification of the semantic settings, structural composition, and narrative context of the content to be generated, rather than simple probabilistic generation based on input data. Accordingly, since the generative AI model participates in the generation process only with data that has passed verification, it is possible to prevent content that is semantically inappropriate or conflicts with the settings from proceeding to the generation stage.

[0027] Furthermore, according to some embodiments of the present disclosure, a spatial and temporal representation including the location and movement of an object is generated only when semantic consistency of the content to be generated is ensured, and an evaluation based on physical constraints can be performed based on this representation. Through this, physical errors such as penetration between objects, unnatural positional changes due to gravity violations, movement ignoring friction conditions, or abrupt changes in movement contrary to inertia can be identified at a stage prior to content generation. In other words, a structure can be provided that allows for the detection and response to physical errors in advance during the generation process, rather than discovering them retrospectively in the resulting product.

[0028] According to some embodiments of the present disclosure, if it is determined that physical constraints are not satisfied, the generation request is not merely discarded; instead, analysis data including the location and extent of the discrepancy may be extracted. This analysis data is converted into control information interpretable by the generative AI model and reflected back into the generation process, thereby preventing the repeated generation of areas or patterns where physical errors occurred. Accordingly, the effect of automatically correcting the generation result can be achieved without changing the internal structure or learning parameters of the generative AI model.

[0029] Consequently, the generative AI content generation control technique according to some embodiments of the present disclosure combines semantic verification and physical verification to control the entire generation process step-by-step, thereby simultaneously ensuring spatiotemporal consistency and physical consistency of generative AI content. This provides the effect of automatically generating stable and reliable high-quality content without relying on random regeneration or post-production editing.

[0030] The effects obtainable through the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which the present disclosure belongs from the description below. Brief explanation of the drawing

[0031] Various embodiments of the present disclosure are described with reference to the drawings, wherein similar reference numbers are used to refer to collectively similar components. In the following embodiments, for illustrative purposes, a number of specific details are presented to provide a comprehensive understanding of one or more embodiments. However, it will be apparent that such embodiment(s) may be practiced without these specific details. FIG. 1 is a block diagram for illustrating an apparatus according to some embodiments of the present disclosure. FIG. 2 is a flowchart illustrating an example of a method for generating final content using a generative artificial intelligence model according to some embodiments of the present disclosure. FIG. 3 is a flowchart illustrating an example of a method for determining whether a first data set satisfies a predefined generation condition according to some embodiments of the present disclosure. FIG. 4 is a flowchart illustrating an example of a method for determining whether the location and operation of an object satisfy predefined constraints according to some embodiments of the present disclosure. FIG. 5 is a flowchart illustrating an example of a method for generating final content by providing control information as a conditional input to a generative artificial intelligence model according to some embodiments of the present disclosure. Specific details for implementing the invention

[0032] Hereinafter, various embodiment(s) of the apparatus and the method for controlling the apparatus according to the present disclosure will be described in detail with reference to the drawings. Identical or similar components are given the same reference numeral regardless of the drawing symbols, and redundant descriptions thereof will be omitted.

[0033] The purpose and effects of the present disclosure, and the technical configurations for achieving them, will become clear by referring to the embodiments described in detail below in conjunction with the accompanying drawings. In describing one or more embodiments of the present disclosure, if it is determined that a detailed description of related prior art may obscure the essence of at least one embodiment of the present disclosure, such detailed description is omitted.

[0034] The terms of this disclosure are defined with consideration of their functions in this disclosure, and these may vary depending on the intentions or practices of the user or operator. Furthermore, the attached drawings are intended only to facilitate understanding of one or more embodiments of this disclosure, and the technical scope of this disclosure is not limited by the attached drawings; it should be understood that the technical scope of this disclosure includes all modifications, equivalents, and substitutions that fall within the spirit and technical scope of the invention.

[0035] The suffixes "module" and "part" for components used in the following description are assigned or used interchangeably solely for the ease of drafting the present disclosure, and do not have distinct meanings or roles in themselves.

[0036] Terms including ordinal numbers, such as first, second, etc., may be used to describe various components, but said components are not limited by said terms. Such terms are used solely for the purpose of distinguishing one component from another. Accordingly, the first component mentioned below may be the second component within the technical scope of the present disclosure.

[0037] Singular expressions include plural expressions unless the context clearly indicates otherwise. That is, unless otherwise specified or the context does not make it clear that the singular form is indicated, the singular in this disclosure and claims should generally be interpreted to mean "one or more."

[0038] In this disclosure, terms such as “comprising,” “comprising,” or “having” are intended to specify the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in this disclosure, and should not be understood as precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0039] In this disclosure, the term “or” should be understood as “or” in an implied sense, not in an exclusive sense. That is, unless otherwise specified or evident from the context, “X uses A or B” is intended to mean one of the natural implied substitutions. That is, where X uses A; where X uses B; or where X uses both A and B, “X uses A or B” may apply to any of these cases. Furthermore, the term “and / or” as used in this disclosure should be understood to refer to and include all possible combinations of one or more of the listed related items.

[0040] The expression 'at least one of A, B, and C' means one or more of the elements of the group consisting of A, B, and C, and should not be interpreted as requiring at least one of each of the listed A, B, and C, regardless of whether A, B, and C are related as categories.

[0041] The terms "information" and "data" as used in this disclosure may be used interchangeably.

[0042] Unless otherwise defined, all terms used in this disclosure (including technical and scientific terms) may be used in a meaning that is commonly understood by those skilled in the art of this disclosure. Additionally, terms defined in commonly used dictionaries are not to be over-interpreted unless specifically defined otherwise.

[0043] However, the present disclosure is not limited to the embodiments disclosed below but may be implemented in various different forms. Some embodiments of the present disclosure are provided merely to fully inform those skilled in the art of the scope of the present disclosure, and the present disclosure is defined only by the scope of the claims. Therefore, such definition must be based on the content throughout the present disclosure.

[0044] The present disclosure relates to a method for controlling content generation using a generative artificial intelligence model, and more specifically, to a technology for automatically controlling and correcting the generation process of generative artificial intelligence content by performing verification of semantic settings, narrative context, and physical constraints in stages during the content generation process, and deriving control information from analysis data generated based on the verification results and providing it as a conditional input to the generative artificial intelligence model.

[0045] According to some embodiments of the present disclosure, physical simulation-based verification is performed only when semantic verification is passed, and if it is determined that physical constraints are not satisfied, analysis data including the location and degree of the discrepancy is extracted, and said analysis data is converted into control information interpretable by a generative artificial intelligence model and can be used for the regeneration or correction of content. Hereinafter, a system for multiple verification and automatic correction of generative artificial intelligence content according to one embodiment of the present disclosure will be described with reference to FIGS. 1 to 5.

[0046] FIG. 1 is a block diagram for illustrating an apparatus according to some embodiments of the present disclosure.

[0047] The device (100) described in the present disclosure may include a device capable of transmitting, receiving, processing, or outputting data, content, service, and application, and may be configured to perform content generation control, verification, and automatic correction functions using a generative artificial intelligence model.

[0048] The device (100) of the present disclosure may be paired or connected with another device or external server via a wired or wireless network, thereby transmitting or receiving input data, predefined generation conditions, predefined constraints, analysis data, control information, or generation result data related to the generation process of generative artificial intelligence content. In this case, the data transmitted or received through the device (100) may be converted into a form suitable for the generation control technique of the present invention before or during the transmission and reception process.

[0049] The device (100) of the present disclosure may include, for example, a fixed device such as a personal computer (PC), a microprocessor-based system, a mainframe computer, a digital processor, a device controller, a network TV, a smart TV, an IPTV, a digital TV, digital signage, etc., and a mobile device such as a personal handheld terminal (PDA), a smartphone, a tablet PC, a laptop, etc. Additionally, the device (100) of the present disclosure may also include a server device such as an application server, a computing server, a database server, a file server, a web server, etc., which performs the execution of a generative artificial intelligence model, verification based on physical simulation, or analysis data processing and control information conversion. However, it is not limited thereto.

[0050] In the present disclosure, when referred to as a device (100), the meaning may refer to at least one of a computer system or computer device, a fixed device, a mobile device, or a server that performs the functions of generating and controlling generative artificial intelligence content depending on the context, and unless specifically limited, it may be used to include all of the above devices.

[0051] According to some embodiments of the present disclosure, the device (100) may further include an orchestrator that controls the calling and conditional input configuration of a generative artificial intelligence model, and a physics engine that performs collision evaluation and computations based on physical constraints for simplified objects. The orchestrator and the physics engine may be separate hardware components, but in one embodiment, they may be implemented as a logic module or a software module executed by a processor (110).

[0052] Referring to FIG. 1, the device (100) may include a processor (110), a storage unit (120), and a communication unit. Since the components illustrated in FIG. 1 are not essential for implementing the device (100), the device (100) described in this disclosure may have more or fewer components than those listed above.

[0053] Each component of the device (100) of the present disclosure may be integrated, added, or omitted according to the specifications of the device (100) actually implemented. That is, as needed, two or more components may be combined into one component, or one component may be subdivided into two or more components. Furthermore, the function performed in each block is intended to explain the embodiments of the present disclosure, and the specific operation or device thereof does not limit the scope of the present invention.

[0054] In addition to operations related to the application, the processor (110) can control the overall operation of the device (100). The processor (110) can perform functions related to the generation control and verification of generative artificial intelligence content by processing signals, data, and information input or output through the components of the device (100) or by executing a program stored in the storage unit (120).

[0055] A processor (110) can read a computer program stored in a storage unit (120) and perform a multi-verification and automatic correction method for generative artificial intelligence content according to some embodiments of the present disclosure. To this end, the processor (110) can generate a first data set including semantic information, narrative context information, and structural characteristic information from input data during the content generation process, and perform a function of determining whether the first data set satisfies predefined generation conditions.

[0056] The processor (110) may be configured to determine whether the location of an object and the operation of an object satisfy the predefined physical constraints by generating a second data set containing information regarding the location of an object or the operation of an object included in the content to be generated, when the semantic verification result determines that the content to be generated satisfies the predefined generation conditions, and by performing a simulation or operation according to predefined physical constraints in a virtual space on the second data set.

[0057] According to some embodiments of the present disclosure, if the processor (110) determines that physical constraints are not satisfied during the physical consistency verification process, it may extract analysis data including the location where the discrepancy occurred and the degree of the discrepancy. The processor (110) may perform regeneration or automatic correction of content by converting the analysis data into control information that can be interpreted by a generative artificial intelligence model and providing it as a conditional input to the generative artificial intelligence model.

[0058] According to some embodiments of the present disclosure, a physics engine may be configured to calculate a correct coordinate value (constraint) that satisfies a physical constraint through an inverse kinematics operation when a physical error is detected, return said coordinate value to an orchestrator, and the orchestrator may be configured to update a second data set or guide data based on said coordinate value.

[0059] The processor (110) may include one or more cores and may include at least one of various computing units such as a central processing unit (CPU), a general purpose graphics processing unit (GPGPU), and a tensor processing unit (TPU). However, it is not limited thereto. Additionally, the processor (110) may be configured as a dual processor or other multiprocessor architecture.

[0060] According to some embodiments of the present disclosure, the processor (110) may operate a combination of a CPU and a GPU or TPU to perform operations related to semantic verification, physical simulation-based verification, analysis data generation, and control information conversion in parallel. For example, the execution of a generative artificial intelligence model and the generation of content based on conditional input may be performed on a GPU or TPU, while the determination of predefined generation conditions, data management, and control logic may be performed on a CPU.

[0061] The processor (110) is connected to the storage unit (120) and the communication unit (130) via a high-speed data bus, and can efficiently process large-capacity content data, physical simulation result data, and analysis data. Through this, the input of the generative artificial intelligence model can be controlled in real-time or near-real-time according to the verification results during the content creation process.

[0062] Additionally, the processor (110) can control the creation of content so that the placement of objects, the development of actions, or the temporal flow are maintained by using verified guide data and control information during the creation process of generative artificial intelligence content, and can be configured to create content with physical consistency and spatiotemporal consistency without repetitive regeneration or user intervention.

[0063] The storage unit (120) can store data and programs to support various functions of the device (100). The storage unit (120) can store a number of applications running on the device (100), commands to control the operation of the device (100), and various data related to the generation control of generative artificial intelligence content. At least some of these applications may be received from an external server via the communication unit (130) and stored, and at least some may be stored in the storage unit (120) from the time of shipment to perform the basic functions of the device (100).

[0064] Applications and data stored in the storage unit (120) may be read and executed by the processor (110) to perform multiple verification and automatic correction functions of generative artificial intelligence content according to some embodiments of the present disclosure.

[0065] According to some embodiments of the present disclosure, the storage unit (120) may store at least one of input data used during the generation process of generative artificial intelligence content, a first data set including semantic information, narrative context information or structural characteristic information, a second data set including information on at least one of the location of an object or the operation of an object, analysis data extracted from a verification result according to physical constraints, control information generated based on the analysis data, and guide data including the control information.

[0066] According to some embodiments of the present disclosure, the storage unit (120) may store predefined generation conditions referenced during the generation process of generative artificial intelligence content. However, it is not limited thereto, and the predefined generation conditions referenced during the generation process of generative artificial intelligence content may be provided to the processor (110) through the communication unit (130) in a state stored on an external server.

[0067] The predefined generation conditions may include at least one of the worldview setting, object attributes, structural characteristics, narrative rules, or constraints of the content to be generated, and the processor (110) may be configured to determine whether the first data set satisfies the predefined generation conditions by referring to the predefined generation conditions.

[0068] According to some embodiments of the present disclosure, the storage unit (120) may store predefined constraints used to verify the physical consistency of generative AI content. However, the present disclosure is not limited thereto, and predefined constraints used to verify the physical consistency of generative AI content may also be provided through the communication unit (130) in a state stored on an external server.

[0069] The predefined constraints may include conditions for evaluating at least one of the possibility of collision between objects, change in position due to gravity, continuity of motion due to inertia, limitation of relative movement due to friction, or limitation of the movement range of objects. The processor (110) may be configured to perform a simulation or operation on a second data set by referring to the stored predefined constraints.

[0070] According to some embodiments of the present disclosure, predefined generation conditions and predefined constraints may be managed separately from each other according to the type of content, scene composition, or object characteristics, and may be stored in at least one of a storage unit (120) or an external server. Meanwhile, the processor (110) may be configured to selectively retrieve and apply predefined generation conditions and constraints corresponding to the content to be generated during the content generation process. However, the present disclosure is not limited thereto.

[0071] Additionally, the storage unit (120) can store parameter information related to the execution of a generative artificial intelligence model, conditional input data, reference image or external characteristic data, and content generation result data, and can support the processor (110) in performing content generation, regeneration, or automatic correction using said data.

[0072] The storage unit (120) can store information of any form generated or determined by the processor (110) and information of any form received through the communication unit (130), and the information may include verification results, intermediate generation results, and correction history that occurred during the content generation process.

[0073] The storage unit (120) may include at least one type of storage medium among flash memory type, hard disk type, SSD type (Solid State Disk type), SSD type (Silicon Disk Drive type), multimedia card micro type, card type memory (e.g., SD or XD memory, etc.), RAM (random access memory; RAM), SRAM (static random access memory), ROM (read-only memory; ROM), EEPROM (electrically erasable programmable read-only memory), PROM (programmable read-only memory), magnetic memory, magnetic disk, and optical disk. The device (100) may also be operated in connection with web storage that performs the storage function of the storage unit (120) on the internet.

[0074] According to some embodiments of the present disclosure, the storage unit (120) may include a high-speed storage device, such as a cache memory or NVRAM (Non-Volatile Random Access Memory), for high-speed data exchange with the processor (110). This enables operations such as verification, analysis data generation, and control information conversion performed during the generation process of generative artificial intelligence content to be performed in real-time or near-real-time.

[0075] The communication unit (130) may include one or more modules that enable wired and wireless communication between the device (100) and a wired or wireless communication system, between the device (100) and another device, or between the device (100) and an external server. Additionally, the communication unit (130) may include one or more communication modules for connecting the device (100) to one or more networks.

[0076] The communication unit (130) is a module for wired or wireless internet access, may be embedded in or externally mounted on the device (100), and may be configured to transmit or receive wired or wireless signals. According to some embodiments of the present disclosure, the generation control of generative artificial intelligence content, the determination of predefined generation conditions, and verification operations based on predefined constraints are performed by a processor (110) inside the device (100), and the communication unit (130) may be configured to provide communication functions to assist the operations as needed.

[0077] According to some embodiments of the present disclosure, the communication unit (130) may not be used when the generative artificial intelligence model is executed within the device (100), and may be selectively used only when at least some of the generative artificial intelligence model or physical simulation operations are performed on an external computing resource. For example, when the generative artificial intelligence model or physical simulation is executed on an external server or in a cloud environment, the communication unit (130) may transmit input data, control information, or guide data generated by the processor (110) to an external server and receive the processing results.

[0078] Additionally, the communication unit (130) may be configured to transmit information regarding a second data set or predefined constraints to an external server and receive the results of the physical simulation or analysis data, only when high-complexity operations during the physical simulation or physical verification operation are performed on an external computing resource. In this case as well, the interpretation of the received result data, the generation of analysis data, and the conversion of control information may be performed by a processor (110) inside the device (100).

[0079] According to some embodiments of the present disclosure, the communication unit (130) may support a distributed structure in which at least some of the steps of semantic verification, physical verification, analysis data generation, and control information conversion performed during the process of generating generative artificial intelligence content are performed on different devices or servers. However, even in such cases, the steps may be managed as a single logically continuous generation control process controlled by the processor (110) of the device (100).

[0080] The communication unit (130) can support wireless communication technologies such as 5G NR (New Radio), millimeter wave (mmWave), Wi-Fi 6, Wi-Fi 6E, and Wi-Fi 7 for high-speed data transmission, and can be designed to be compatible with next-generation mobile communication technologies to be developed in the future. Through this, the generation and automatic correction process of generative artificial intelligence content can be performed in real-time or near-real-time, even when linked with external computing resources.

[0081] In addition, the communication unit (130) can support a multi-connectivity function, and by selectively utilizing a communication path according to the network environment, it can be configured to ensure the stability of data transmission even when external computing resources are used.

[0082] FIG. 2 is a flowchart illustrating an example of a method for generating final content using a generative artificial intelligence model according to some embodiments of the present disclosure. FIG. 3 is a flowchart illustrating an example of a method for determining whether a first data set satisfies a predefined generation condition according to some embodiments of the present disclosure. FIG. 4 is a flowchart illustrating an example of a method for determining whether the location of an object and the behavior of an object satisfy predefined constraints according to some embodiments of the present disclosure. FIG. 5 is a flowchart illustrating an example of a method for generating final content by providing control information as a conditional input to a generative artificial intelligence model according to some embodiments of the present disclosure.

[0083] First, the processor (110) can receive input data for content creation from the user.

[0084] In the present disclosure, a user is an entity that requests the creation of content using a generative artificial intelligence model, and may include at least one of an individual user, a content creator, a corporate user, or a system administrator. For example, the user may be an individual creator requesting the creation of images, videos, animations, virtual environments, game content, or story-based content, or may be a company or organization that creates or manages advertisements, video production, metaverse content, educational content, entertainment content, etc. However, the definition of a user is not limited thereto, and an automated system or another artificial intelligence system may also be considered a user when requesting the creation of content to the device (100) according to the present invention.

[0085] In the present disclosure, content is a result generated by a generative artificial intelligence model and may include at least one of text content, image content, video content, audio content, or multimodal content consisting of a combination thereof. For example, the content may be a single image, video including multiple frames, story video composed of scenes, animation including characters, or virtual environment content including the actions and interactions of objects over time. Additionally, the content may be content generated based on a specific worldview, narrative structure, character settings, or physical environment.

[0086] Input data is data provided by a user or automatically generated by a system for the creation of the above content, and may be multimodal data including text data, image data, voice data, video data, or a combination thereof. For example, the input data may be a text prompt including the theme, genre, mood, emotion, appearing objects, character characteristics, scene description, or action instructions of the content. Additionally, the input data may be a reference image, character appearance data, background image, or a part of previously created content to be applied to the content to be created.

[0087] According to some embodiments of the present disclosure, input data may be provided as a single data type or in a combined form of multiple data types. For example, a user may provide one or more reference images along with a text prompt, or input voice data or video data along with a text description to enable a generative AI model to understand a richer context. Such multimodal input data may be used to define the semantic setting, narrative context, and structural characteristics of the content to be generated.

[0088] Meanwhile, the processor (110) can receive input data in various ways. For example, the processor (110) can receive input data through a user interface included in the device (100), and the user interface may include at least one of a keyboard, a mouse, a touch input device, a voice input device, or a graphical user interface (GUI). Additionally, input data may be provided through a file upload method, a drag-and-drop method, an input method through a text input window, or a voice recognition-based input method.

[0089] According to some embodiments of the present disclosure, input data may be received from an external device or an external server through a communication unit (130). For example, a user may request the creation of content from an external server through a web-based interface or application, and the external server may transmit input data to the device (100). Additionally, input data may be transmitted from another generative artificial intelligence system, a content management system, or an automated pipeline, and in such cases, the processor (110) may process the input data as input data received from the user.

[0090] Meanwhile, referring to FIG. 2, when the processor (110) receives input data for content creation from a user, it can process the input data to create a first data set including at least one of semantic information, structural characteristic information, or narrative context information regarding the content to be created (S100).

[0091] Generally, generative AI models generate content by probabilistically interpreting input data provided by users. However, because this approach does not explicitly separate and manage the semantic settings, structural constraints, or narrative contexts contained in the input data, problems such as physical errors, structural inconsistencies, or the collapse of the story progression can frequently occur during the content generation process.

[0092] Accordingly, the present disclosure may include an input data processing step for extracting the core meaning and structure of the content to be generated from the input data, rather than a method of transmitting the input data as is to a generative artificial intelligence model, and converting it into a data form that can be utilized in a subsequent step (S200) for determining whether a predefined generation condition is satisfied.

[0093] According to some embodiments of the present disclosure, when input data is received, the processor (110) may perform an appropriate input data processing process according to the type, quality, composition, and purpose of the input data. The input data processing process may include a series of operations to analyze the meaning and structure of the input data to extract information necessary for verification, and to normalize and structure it into a form that can be compared and evaluated in subsequent steps.

[0094] For example, the processor (110) can determine whether the input data contains only text, contains text and images together, or combines images and metadata, and can call different processing modules for each modality. Additionally, if part of the input data exists on an external server, the processor (110) can receive or download the data through the communication unit (130) and transfer the received data to an internal processing pipeline.

[0095] The input data processing method can be explained in more detail as follows.

[0096] The processor (110) can first perform a normalization process of the input data.

[0097] For example, if the input data includes text data, the processor (110) can perform the process of unifying the character encoding for the text data, performing sentence separation and tokenization, and standardizing expressions of the same meaning. For instance, if the same object is repeated in different expressions in user input (e.g., "boy," "child," "male child"), it can be normalized into a single standard entity. Additionally, normalization rules or language model-based corrections can be performed so that semantic interpretation is possible even when there is language mixing, typos, or abbreviated expressions.

[0098] In another example, the processor (110) can normalize the resolution, color space, format, and orientation information of the image data when the input data includes image data, and normalize the frame rate, resolution, time code, and compression characteristics of the image data when the input data includes video data.

[0099] The examples described above are merely illustrative examples for explaining the present disclosure, and the present disclosure is not limited to the examples described above.

[0100] The aforementioned normalization method is a preparatory process for reliably applying the same verification criteria even when the sources of input data are diverse, and can reduce errors caused by representation instability or format differences when compared with predefined generation conditions and constraints in subsequent stages.

[0101] After normalization, the processor (110) can perform parsing and segmentation to break down the input data into semantic units.

[0102] For example, if the input data includes text data, the processor (110) can break down sentences within the text data into clauses or phrases and identify noun phrases, verb phrases, modifier phrases, etc., to derive object candidates, action candidates, and attribute candidates. For instance, in the sentence “A person wearing a coat runs down an alley holding an umbrella,” object candidates such as “person,” “coat,” “umbrella,” and “alley,” and action candidates such as “wear,” “hold,” and “run” can be extracted, and modifier structures such as “wearing a coat” can be used as information implying the relationship between the person and the coat and the external attributes of the person.

[0103] In another example, when the input data includes image data, the processor (110) can separate people, objects, and background elements from the image data through object detection and derive the location, size, shape, and features of each object.

[0104] As another example, when the input data includes video data, the processor (110) can separate the sequence of video data into scenes by detecting scene transitions along with frame-by-frame analysis, and can detect changes in the movement or interaction of objects within each scene.

[0105] The examples described above are merely illustrative of the present disclosure and the present disclosure is not limited to the examples described above.

[0106] The aforementioned decomposition process is a step of separating and extracting complex information within input data into components necessary for verification, and can contribute to subsequent information definition and data set generation.

[0107] Meanwhile, the processor (110) can perform feature extraction on the disassembled components.

[0108] For example, if the input data includes text data, the processor (110) can extract keywords, named entities, attribute expressions, relationship expressions, time expressions, condition expressions, prohibition or necessity expressions from the text data, and can identify the narrative tone through emotion or mood expressions.

[0109] As another example, when the input data includes image data, the processor (110) can extract features regarding structural characteristics and external characteristics, such as feature points, contours, texture features, color tone, lighting atmosphere, approximate pose, and viewing direction, from the image data.

[0110] As another example, when the input data includes image data, the processor (110) can estimate the movement path and speed change of the object included in the image data, whether there is contact between the objects included in the image data and the timing of the interaction, and the continuity between scenes from the image data through object tracking, and can extract the state change over time.

[0111] The aforementioned features are not merely intended to be provided to a generative AI model, but can also be used to construct a validation input to determine whether the content to be generated matches predefined generation conditions.

[0112] Meanwhile, according to some embodiments of the present disclosure, when input data is multimodal, the processor (110) can perform matching between modalities. For example, if an attribute such as "red coat" mentioned in text corresponds to the color of a person's clothing detected in a reference image, the object / attribute information extracted from the text and the feature extracted from the image can be mapped to the same entity. Additionally, it can check whether an instruction such as "move to the right" in the text matches the direction of movement appearing in the image or sketch, or indicate a contradiction if they are mutually contradictory. This integration process contributes to converging information expressed by different modalities into a single definition of content to be generated, and can reduce the problem of different representations conflicting for the same object in a subsequent verification step. The integration process can be performed through rule-based matching, embedding-based similarity matching, or comparison with predefined reference data, but is not limited thereto.

[0113] Meanwhile, in the present disclosure, the processor (110) can generate at least one of semantic information, structural characteristic information, and narrative context information regarding the content to be generated based on the result of processing input data.

[0114] Semantic information is information that represents the semantic level of the content to be generated, and may include semantic relationships between objects or characters, backgrounds, events, attributes, relationships, actions, and concepts. For example, the semantic information "a person is holding an umbrella" may include the entity "person," the entity "umbrella," the action "holding," and the relationship between the person and the umbrella. Semantic information may also include semantic constraints related to the worldview or character settings. For instance, if the input contains an action that a specific character cannot perform within a particular worldview, this part can be extracted as semantic information and compared with predefined generation conditions in a subsequent step. Semantic information can be organized into text-based representations, entity-relationship graphs, or vectorizable representations, and this may vary depending on the implementation.

[0115] Structural characteristic information is information that expresses the composition and development structure of the content to be generated. Structural characteristic information may include scene divisions, time interval divisions, relative placement relationships of objects, stepwise composition of actions, camera viewpoints, or screen compositions. For example, a structural characteristic such as "a character is located on the left side of the screen in Scene 1 and moves to the right in Scene 2" can be viewed as structural information that includes changes in the relative position of objects and transitions between time intervals. Structural characteristic information can serve as a basis for generating a second data set in a subsequent step and, in particular, can provide the minimum spatial and temporal framework necessary to perform physical verification. In this disclosure, structural characteristics sufficient for physical simulation or constraint evaluation can be extracted, rather than detailed structures at the level of high-resolution rendering.

[0116] Narrative contextual information refers to contextual information that ensures the content being generated possesses narrative continuity rather than being a mere sequence of scenes. Narrative contextual information can include causal relationships between events, the temporal order of events, character goals and motivations, changes in state, and conditions for continuity between scenes. For example, the input "The character runs into an alley to escape pursuit" includes not only the action of simply "running" but also the motivation of "escaping pursuit" and a cause-and-effect connection, which can be extracted as narrative contextual information. Furthermore, continuity conditions such as "If the character was holding an umbrella in the previous scene, the umbrella must remain in the next scene" can be included as narrative contextual information and can be utilized to determine consistency with setting data during subsequent semantic verification.

[0117] Meanwhile, the processor (110) may generate a first data set to include at least one of the semantic information, structural characteristic information, and narrative context information. The first data set is a data set structured to suit the purpose of verifying the input data processing results, and may be designed to enable comparison with predefined generation conditions in a subsequent step and further lead to a physical verification step.

[0118] The first dataset is not limited to a single format and may consist of an entity list, attribute records, relationship records, event records, time interval records, structure records, or a combination thereof. Additionally, the first dataset may include, but is not limited to, a text-based normalized representation or an embeddingable representation to enable vector space comparison to be performed in a subsequent step. For example, the first dataset may be stored in the form of a graph having an "entity-relationship-event" structure, where the nodes of the graph represent entities or events, and the edges represent relationships or causations. Alternatively, the first dataset may include an entity-specific attribute set and a relationship set as a key-value structure. Additionally, depending on the implementation, the first dataset may simultaneously include a structure close to the raw data (e.g., a list of extracted keywords and relationships) and a representation for verification (e.g., an embedding vector).

[0119] The reason a first dataset is required is that, in order to perform verification-based generation control in this disclosure, input data must be transformed into a verifiable structure rather than being used as is. When input data is provided in natural language, the same meaning may appear in various expressions, some information may be implicitly included, and inconsistencies in expression between modalities may occur. In such a state, it is difficult to reliably determine consistency with predefined generation conditions, and structural premises for evaluating physical constraints may become unclear. The first dataset extracts key elements necessary for verification from the input data, normalizes them, creates comparable expressions, and can function as an intermediate expression connecting the pipeline leading to semantic verification and physical verification.

[0120] Furthermore, the first data set can also contribute to the computational efficiency of subsequent steps. For example, if predefined generation conditions are not met, it is possible to control the process so that physical simulations or control information conversions in subsequent steps are not performed; to achieve this, information required for comparison with predefined generation conditions must be prepared in advance in the form of the first data set.

[0121] To explain the process of generating the first data set in more detail, the processor (110) can extract entity candidates from the input data processing results and standardize them. Entity candidates can be derived from named entities and noun phrases in text, object detection results in images / videos, or asset lists explicitly provided by the user. Subsequently, the processor (110) can extract entity-specific attribute candidates, which can be derived from formula expressions in text, emotion / personality expressions, or image feature extraction results. Next, the processor (110) can construct relationship and event information between entities. Relationships can be expressed in various forms, such as location relationships, interaction relationships, ownership relationships, usage relationships, contact relationships, etc., and events can indicate "what happens when and how" by including one or more relationships and temporal markers. Additionally, the processor (110) can add causal relationships or state transitions between events by analyzing cause-effect expressions, goal-action connections, state change expressions, etc., to construct a narrative context. From the perspective of structural characteristics, events and relationships can be mapped to specific scenes or specific time intervals using the results of scene division or time interval division, and a step structure such as the start-progress-end of an action can be formed. Finally, the processor (110) can reconstruct the first data set into a normalized text representation or a vectorizable representation so that vector space comparison is possible in a subsequent step. For example, entities, attributes, relationships, and events can be serialized into a text sequence using a certain template, or an entity-relationship graph can be converted into an embedding input. Such representation generation can be stored as part of the first data set and can be used to calculate similarity with predefined generation conditions in a subsequent step.

[0122] Meanwhile, the processor (110) can determine whether the first data set satisfies the predefined generation conditions by comparing the first data set generated in step (S100) and the content to be generated in a vector space (S200).

[0123] In the present disclosure, predefined generation conditions refer to semantic and logical standards that the content to be generated must comply with, and can be understood as a data set that defines the worldview, character settings, narrative rules, prohibitions, or requirements of the content to be generated by the generative AI model. Predefined generation conditions are managed independently of input data and may be explicitly set by the user, or may be predefined by a system administrator, content creator, or service provider and stored in the storage unit (120). Additionally, predefined generation conditions may be configured differently depending on the type, genre, service purpose, or platform policy of the content.

[0124] For example, predefined generation conditions may include rules that restrict the generation of events or actions not permitted in a specific worldview within content based on that worldview.

[0125] According to some embodiments of the present disclosure, a condition that modern firearms or automobiles must not appear in a medieval fantasy worldview may be defined as a generation condition, and a condition that unscientific explanations or false information are not generated in educational content based on scientific facts may be included.

[0126] According to some other embodiments of the present disclosure, if the personality of a specific character is predefined, a condition to prevent the character from acting contrary to that personality may be included in the generation condition. For example, a scene in which a character defined as "introverted and non-violent" randomly performs aggressive behavior may be judged as a violation of the generation condition.

[0127] Furthermore, predefined generation conditions may include conditions for maintaining narrative continuity. For example, generation conditions can be defined as a continuity condition requiring a character possessing a specific object in a previous scene to possess the same object in the next scene, a causal condition requiring that a subsequent event can only occur after a specific event has taken place, or a condition requiring that a specific action cannot be performed before a specific state change occurs. These conditions can be particularly important when generating content that spans multiple scenes or time intervals, rather than just a single scene.

[0128] Furthermore, predefined creation conditions may include the level of expression of the content, ethical standards, and policy criteria. For example, conditions may include requirements that violent expressions, hate speech, or inappropriate interactions be restricted in content provided to specific age groups, or conditions may be included to prevent the creation of elements that may infringe copyright in accordance with specific service policies.

[0129] According to some embodiments of the present disclosure, predefined generation conditions are not limited to a single criterion but may include a combination of semantic, logical, narrative, and policy criteria.

[0130] In step (S200), the processor (110) may determine consistency through comparison in a vector space rather than determining the first data set and predefined generation conditions by a direct rule comparison method. This is because input data and generation conditions are often expressed as unstructured data such as natural language, image descriptions, and narrative rules, making it difficult to reliably determine semantic similarity or inconsistency using only simple string comparison or rule-based matching. Comparison in a vector space allows for numerical expression of semantic similarity, conceptual distance, and contextual difference, thereby enabling determination of whether there is semantic agreement or inconsistency despite various methods of expression.

[0131] Specifically, referring to FIG. 3, the processor (110) can vectorize the first data set and predefined generation conditions, respectively, and represent them in a vector space (S210). Here, vectorization refers to the process of converting text, graphs, relational structures, or other unstructured data into numerical vectors of fixed dimensions, and can be configured so that semantic similarity is reflected as distance or angle between vectors. For example, semantic information, structural characteristic information, or narrative context information included in the first data set can be serialized into text sequences, relational representations, or event representations, and then converted into vectors through a pre-trained embedding model.

[0132] Predefined generation conditions can also be vectorized in the same or compatible manner. For example, if world setting documents, character setting phrases, narrative rule sentences, or policy rules are stored in text form, the processor (110) can convert them into vectors using the same embedding model as the first data set. In this case, by mapping the first data set and the predefined generation conditions into the same vector space, the semantic similarity or difference between the two data can be quantitatively compared.

[0133] In this disclosure, a vector space refers to an abstract semantic space in which vectorized data is located, configured such that semantically similar concepts or contexts are placed close to each other, while semantically different or conflicting concepts are placed far apart. Such a vector space may be a high-dimensional space, and each dimension may reflect specific semantic features or latent concepts. The number of dimensions, the configuration method, or the learning method of the vector space may vary depending on the implementation, and this disclosure is not limited to a specific embedding technique or number of dimensions.

[0134] Meanwhile, the processor (110) can calculate the similarity between a vector corresponding to a first data set and a vector corresponding to a predefined generation condition. As a method for calculating similarity, cosine similarity, inner product-based similarity, Euclidean distance-based similarity, or other statistical or geometric distance measures may be used, but are not limited thereto. For example, cosine similarity can enable a comparison that prioritizes directionality over the magnitude of vectors by determining similarity based on the angle between two vectors.

[0135] The result of calculating similarity can be expressed as a numerical value, which indicates how semantically similar or consistent the first dataset is with the predefined generation conditions. For example, a higher similarity value may indicate that the semantic and narrative information contained in the first dataset aligns well with the predefined generation conditions, while a lower similarity value may indicate a higher probability of discrepancy with the predefined generation conditions.

[0136] Meanwhile, the processor (110) can determine whether a predefined generation condition is satisfied based on whether the calculated similarity satisfies a reference value (S220). Here, the reference value is a threshold value pre-set to determine whether the predefined generation condition is satisfied, and can be defined differently depending on the service purpose, content type, the strictness of verification, or user settings. For example, a high reference value may be set for content requiring strict worldview consistency, and a relatively low reference value may be set for creative content with a high degree of freedom.

[0137] The reference value may be a single fixed value or may consist of multiple reference values.

[0138] According to some embodiments of the present disclosure, the processor (110) may be configured to determine that a predefined generation condition is satisfied when the similarity value is greater than or equal to a first reference value, and to determine that a predefined generation condition is not satisfied when the similarity value is less than or equal to a second reference value which is smaller than the first reference value, and to perform additional auxiliary judgment in an intermediate interval between the first reference value and the second reference value.

[0139] According to some other embodiments of the present disclosure, the processor (110) may apply different reference values ​​for different generation condition items. For example, the processor (110) may apply a high reference value to generation conditions related to character personality and a relatively low reference value to generation conditions related to background setting.

[0140] In the present disclosure, the statement that similarity satisfies a reference value may mean a state in which the similarity value between a first data set calculated in a vector space and a predefined generation condition satisfies a predetermined judgment criterion for determining whether the predefined generation condition is satisfied. Here, the reference value may be defined as a single threshold value, or as an allowable range or judgment interval composed of one or more values, and the method of determining whether similarity satisfies the reference value may be implemented in various ways according to embodiments of the present disclosure.

[0141] According to some embodiments of the present disclosure, satisfying a reference value for similarity may mean that the calculated similarity value is greater than or equal to the reference value. For example, if the similarity value is normalized to a range of 0 to 1, and the reference value is set to 0.8, it may be determined that the predefined generation condition is satisfied when the similarity between the first data set and the predefined generation condition is 0.8 or greater. In such embodiments, the predefined generation condition is determined to be satisfied only when the similarity exceeds the reference value or is greater than or equal to the reference value, and it may be determined that the predefined generation condition is not satisfied when the similarity is less than the reference value. This method may be suitable for generating content that requires a high level of semantic consistency, such as maintaining worldview consistency or character personality.

[0142] According to some other embodiments of the present disclosure, satisfying a reference value for similarity may mean that the similarity value falls within a predefined allowable range. For example, a case where the similarity value falls within a certain range, such as 0.6 or higher and 0.9 or lower, may be determined as satisfying a predefined generation condition. In such embodiments, considering cases where an excessively high similarity value may limit the diversity of expression, an allowable range including both an upper and lower limit may be established. For example, since very high similarity may lead to results that repeat or mimic the predefined generation condition exactly, a case where the similarity within a certain range is satisfied may be determined as a state of satisfying a suitable generation condition.

[0143] According to some other embodiments of the present disclosure, the reference value may be defined as a plurality of intervals rather than a single value. For example, if the similarity value is greater than or equal to a first reference value, it may be determined that the predefined generation conditions are fully satisfied; if the similarity value is greater than or equal to a second reference value but less than the first reference value, it may be determined that the predefined generation conditions are partially satisfied; and if the similarity value is less than the second reference value, it may be determined that the predefined generation conditions are not satisfied. In such a multi-stage determination method, control may be exercised so that a subsequent step is performed only when the predefined generation conditions are fully satisfied, or control may be exercised so that an auxiliary adjustment or additional verification step is performed when the conditions are partially satisfied.

[0144] According to some other embodiments of the present disclosure, a reference value may be defined as a relative comparison standard rather than an absolute value. For example, a first data set may be determined to be most similar to a generation condition among a plurality of generation condition candidates, and the generation condition may be determined to be satisfied if the similarity with the generation condition having the highest similarity among them exceeds a certain percentage. In this case, even if the similarity value does not exceed a specific absolute standard, the generation condition may be determined to be satisfied if it is sufficiently close to the generation condition showing the highest relative consistency. This method may be useful when there are multiple generation conditions or when they need to be dynamically selected based on the genre or context of the content.

[0145] According to some other embodiments of the present disclosure, the reference value may be set differently for each item of the predefined generation conditions. For example, a high reference value may be applied to generation condition items that cause fatal errors upon violation, such as character personality or world setting, and a low reference value may be applied to relatively flexible items, such as background description or atmosphere. In this case, a composite result may be obtained in which some of the information included in the first data set satisfies the reference value and others do not, and the processor (110) may synthesize these item-by-item judgment results to finally determine whether the predefined generation conditions are satisfied.

[0146] According to some other embodiments of the present disclosure, satisfying a reference value for similarity may mean a state in which there are no elements violating a predefined generation condition. For example, if a predefined generation condition is defined as "a specific character does not engage in violent behavior," and the first data set does not contain semantic information implying violent behavior, the similarity value may be determined to fall within the reference range. In this case, the similarity determination may be performed based on the presence or absence of a violating element, and the reference value may be utilized as a threshold value that numerically expresses the probability of violation.

[0147] As described above, in this disclosure, the concept that similarity satisfies a threshold value is not merely limited to "exceeding the threshold value" or "greater than or equal to the threshold value," but may comprehensively refer to a state in which one or more criteria are satisfied to determine semantic consistency, logical compatibility, and violation between a predefined generation condition and a first data set. Accordingly, the method of setting the threshold value, the method of determining similarity, and the logic for determining satisfaction may be implemented in various ways according to several embodiments of this disclosure, and this disclosure is not limited to the definition of a specific threshold value or a specific method of determination.

[0148] Meanwhile, referring again to FIG. 2, if the processor (110) determines that the first data set does not satisfy predefined generation conditions (S200, No), it can control that subsequent steps for generating final content are not performed.

[0149] Specifically, the processor (110) may determine that the predefined generation conditions are not satisfied if the similarity does not meet a reference value based on the result of a comparison in a vector space between the first data set and the predefined generation conditions. In this case, if it is determined that the predefined generation conditions are not satisfied, the processor (110) may control the subsequent steps, such as the second data set generation step, the evaluation step based on physical constraints, the analysis data extraction step, the control information conversion step, or the final content generation step, not to be performed. That is, if the predefined generation conditions are not satisfied, all generation-related operations after semantic verification may be omitted or stopped.

[0150] This step is a control step designed to prevent semantically inappropriate content from being generated by a generative AI model, and can be defined as a step that stops the generation process without consuming unnecessary computational resources even when the semantic settings or narrative context included in the input data do not match predefined generation conditions.

[0151] According to some embodiments of the present disclosure, if it is determined that a predefined generation condition is not satisfied, the processor (110) may immediately terminate the entire generation process or switch to another input data processing path that is likely to satisfy the predefined generation condition. For example, if input data provided by a user clearly conflicts with a specific worldview setting, the processor (110) may reject a generation request based on said input data, provide the user with information regarding the violation of the predefined generation condition, or guide the user to modifyable input elements. However, the control according to some embodiments of the present disclosure may be configured to stop the generation process without requiring user intervention and without providing separate feedback to the user.

[0152] According to some other embodiments of the present disclosure, even when it is determined that a predefined generation condition is not satisfied, the processor (110) may not simply stop all subsequent steps, but may selectively control whether to perform subsequent steps depending on the degree of violation of the predefined generation condition. For example, if the violation of the predefined generation condition is minor, the processor may modify a portion of the input data or reconstruct the first data set before the subsequent physical verification step to determine again whether the predefined generation condition is satisfied. On the other hand, if the violation of the predefined generation condition is significant, the processor may control the termination of the generation request without performing subsequent steps.

[0153] According to some embodiments of the present disclosure, if it is determined that a predefined generation condition is not satisfied, the processor (110) may control the call to the generative artificial intelligence model itself not to be performed. This is to prevent the generation process involving high-cost computation from being performed unnecessarily by blocking the generation request at a stage prior to the execution of the generative artificial intelligence model. For example, if the generative artificial intelligence model is composed of a large-scale neural network model and requires significant computational resources, if it is determined at the semantic verification stage that a predefined generation condition is not satisfied, the generation process may be terminated without calling the model.

[0154] According to some other embodiments of the present disclosure, if it is determined that a predefined generation condition is not satisfied, the processor (110) may discard intermediate data that has already been prepared or calculated during the generation process, or invalidate it so that it is not referenced in subsequent steps. For example, if intermediate representations or temporary data generated during the generation of a first data set exist in the storage unit (120), the processor (110) may set a flag to prevent such data from being used in subsequent steps, or delete the data. This prevents data generated in a state of violation of the predefined generation condition from unintentionally affecting subsequent generation processes.

[0155] According to some other embodiments of the present disclosure, if the processor (110) determines that a predefined generation condition is not satisfied, it may update state information corresponding to the generation request to indicate that the generation request is in a "condition not satisfied" state. This state information may be stored in a storage unit (120), and if the same or similar generation request is repeated thereafter, the processor (110) may be configured to refer to it to efficiently perform re-verification of the same condition violation. However, this state management method is an optional embodiment and does not limit the present disclosure thereto.

[0156] Additionally, control over cases where predefined generation conditions are not met may not be limited to completely halting the entire generation process. According to some embodiments of the present disclosure, if the processor (110) determines that predefined generation conditions are not met, it may be configured to temporarily suspend the execution of subsequent steps, automatically adjust a portion of the input data so that the predefined generation conditions can be met, or apply alternative setting data and then determine again whether the predefined generation conditions are met. However, whether such retries or corrections are performed is also controlled by the processor (110), and the common feature is that random generation is not allowed to proceed while the predefined generation conditions are not met.

[0157] Accordingly, the step of controlling subsequent steps so that they are not performed when it is determined that the first data set does not satisfy predefined generation conditions can serve to ensure that the generation process of generative AI content proceeds only when it passes semantic verification. Through this, the generative AI content generation control method according to some embodiments of the present disclosure can implement a stepwise generation control structure in which physical verification and final generation are performed only when semantic consistency is secured, without relying on random generation or post-editing.

[0158] On the other hand, if the processor (110) determines that the first data set satisfies the generation conditions (S200, Yes), it may generate a second data set based on the first data set that includes information on at least one of the location of an object or the operation of an object included in the content to be generated (S300).

[0159] In the present disclosure, the second data set may refer to a data set comprising at least one of position, orientation, size, pose, movement path, velocity, acceleration, contact state, or interaction state at at least one point in time or time interval for one or more objects included in the content to be generated. The second data set does not require a complete 3D model for high-resolution rendering or a precise physical model, and the key aspect is to approximate the spatial arrangement and behavior of the objects to a level sufficient to allow for evaluation by placing simplified objects and applying physical constraints in subsequent steps. Accordingly, the second data set is not limited to accurately restoring the position or behavior of the objects, but may include being configured in a form that enables physical verification within the scope of satisfying the structural constraints and semantic conditions of the content to be generated.

[0160] According to some embodiments of the present disclosure, the processor (110) may determine a set of objects to be included in a second data set by using entity (object / character / background) definitions, attribute definitions, relationship definitions, and event definitions included in a first data set. For example, if a relationship / event such as "Character A grabs a cup on a table" is defined in the first data set, the processor (110) may include Character A, a table, and a cup as a set of objects in the second data set, and additionally include a state indicating a contact relationship between a hand (or arm) and a cup to evaluate the interaction of "grabbing." Additionally, if context information such as "two people face each other" is included in the first data set, the processor (110) may be configured to include motion-related information, such as the relative positions and gaze directions of the two people, in the second data set.

[0161] The processor (110) can define the time axis configuration of the second data set using structural characteristic information and narrative context information of the first data set. For example, if the content to be generated is a single scene, the second data set may be configured as a single point in time or a single time interval, and if it is configured as multiple scenes or multiple time intervals, object states corresponding to each scene or each time interval may be included in the second data set in the form of a sequence. In this case, the length of the time interval, the sampling interval, or the state update cycle may vary depending on the implementation and do not necessarily have to be a constant interval. For example, denser time sampling may be applied in intervals with a high probability of physical error, such as event transition points or contact occurrence points, and coarser sampling may be applied in simple movement intervals.

[0162] To describe a method for generating information about the location of an object, the processor (110) can determine an initial or target location for each object by using relative positional relationships included in a first data set (e.g., "A is on top of B," "A is to the left of B," "A is in front of B"), scene composition information (e.g., "The person is in the center of the screen," "The table is in the foreground," "The background is in the background"), or spatial constraint information (e.g., "The door is attached to the wall," "The object is placed on the floor"). The location may be expressed in 2D coordinates or 3D coordinates. For example, in the case of image generation or 2D-based content, the location may be expressed as (x, y) in the screen coordinate system or image coordinate system, and in the case of constructing a virtual space for physical simulation-based verification, the location may be expressed as (x, y, z) in the world coordinate system. Additionally, while the location can be directly specified using absolute coordinates, it can also be expressed using relative coordinates between objects (e.g., the cup is located at +0.2m and +0.1m from the center of the table).

[0163] According to some embodiments of the present disclosure, the processor (110) may set an "anchor object" or a "reference coordinate system" and then place other objects relative to it. For example, the floor, table, or background structure of a scene may be set as the anchor object, and objects such as a cup, a book, or a person's hand may be placed relative to the anchor object. This method of placement is advantageous when evaluating the direction of gravity, contact surfaces, and collision probability in a physics simulation, and can be naturally connected to the process of placing simplified objects in a virtual space in a subsequent step. Additionally, if the size, shape, or material properties of the objects (e.g., a large box, a small cup, a long rod) are included in the first data set, the processor (110) may perform placement considering the spacing or contact surfaces between objects by reflecting these properties. For example, if the event "a person sits on a chair" is included, the position of the person may be set so that the height of the chair seat and the position of the person's buttocks match, and if the event "a cup is placed on a table" is included, the position of the cup may be set so that the bottom surface of the cup contacts the table top.

[0164] Additionally, the processor (110) may define a change in position using temporal sequence or event development information included in the first data set. For example, if structural characteristic information such as “a person moves from left to right” is included, the processor (110) may define a starting position and an ending position and include a movement path between them in the second data set. The movement path may be defined as a linear path, a curved path, or a path that avoids obstacles, and the path may vary according to relationship information included in the first data set (e.g., “moving while avoiding walls”). In this case, the position for each time interval may be defined as a sampling point on the path, and if a change in speed (e.g., “walking slowly and then starting to run”) is specified, the amount of change in position for each time interval may be set to vary.

[0165] To describe a method for generating information regarding the motion of an object, the processor (110) can define the motion state of each object using event information and narrative context information included in the first data set. Here, the motion information may not be simply binary information such as "moves / stops," but may include the object's pose, orientation, joint state, contact state, interaction state, movement speed and acceleration, or motion steps (e.g., reaching out, grabbing, then lifting). For example, if the event "a person opens a door" is included, the motion information may be configured to include the motion of the person's arm joint moving to the door handle, the contact state of the hand grasping the handle, the state of the handle rotating, and the state of the door opening around the axis of rotation. Of course, such motion information does not need to include all joint angles at the level of high-precision animation, but may be included as an approximated representation at a level where physical constraints can be evaluated.

[0166] According to some embodiments of the present disclosure, the processor (110) may use a motion template or a behavior primitive to generate motion information. For example, if motion types such as "walking," "running," "grabbing," "pushing," "throwing," "sitting," "standing," and "opening a door" are predefined, and a verb or action expression included in an event record of a first data set is mapped to one of these motion types, the processor (110) may call the corresponding motion template to set the step structure and key parameters of the motion. The parameters may include start / end positions, target objects, contact points, movement speed, duration, etc., and if attribute information included in the first data set (e.g., "quickly," "carefully," "with all one's might") exists, it may be reflected in the corresponding parameters.

[0167] Additionally, motion information can be specified by relationship information between objects. For example, in the event "a person picks up a cup," if the position of the cup is included in a second dataset, the target position of the person's hand (or simplified upper limb object) can be automatically set based on the position of the cup. If the event "two objects collide" is included, the collision time, direction, and velocity can be included as motion information. If the event "a person pushes and moves a table" is included, conditions for maintaining contact between the person and the table, the direction of the table's movement, and the velocity considering friction can be included as motion information. In this way, by configuring motion information to include not only the state change of a single object but also interactions between objects, it becomes easier to evaluate whether constraints are satisfied in subsequent steps.

[0168] According to some embodiments of the present disclosure, the processor (110) may be configured to include multiple candidate locations or candidate movements when generating a second data set, so that information regarding the location and movement of an object is not calculated as an "absolutely determined correct answer," but rather one or more candidate results can be calculated in a subsequent physical evaluation step. For example, if an event such as "a person steps over a desk" is included, there may be multiple possible movement trajectories depending on the height of the desk or the length of the person's legs, and the processor (110) may include multiple candidate trajectories in the second data set and then select a suitable candidate according to constraints in a subsequent step. Such a candidate-based configuration can be naturally combined with a structure that adopts only the matched result by applying physical constraints in a subsequent step.

[0169] Consequently, in step (S300), the processor (110) can generate a second data set that approximates the location and / or behavior of an object appearing in the content to be generated, based on at least one of semantic information, structural characteristic information, or narrative context information included in the first data set.

[0170] Meanwhile, if the processor (110) has generated a second data set in step (S300), it can evaluate whether the second data set satisfies predefined constraints (S400). That is, the processor (110) can determine whether the location of an object and the behavior of an object satisfy predefined constraints by performing an evaluation on the second data set according to predefined constraints.

[0171] In the present disclosure, a predefined constraint refers to an evaluation criterion for determining whether the location and behavior of an object included in a second data set exist within an acceptable range, and may include conditions regarding at least one of physical laws, environmental rules, or rules of interaction between objects. A predefined constraint may consist of a single condition, or it may consist of a set of conditions in which multiple conditions are combined. Additionally, predefined constraints may be set differently depending on the content type (e.g., realistic video, animation, game scene), scene characteristics (e.g., indoor / outdoor, scene with gravity direction, underwater scene), and object characteristics (e.g., rigid body, soft body, stationary object, moving object), and may be stored in a storage unit (120) and then called and applied by a processor (110). For example, in content requiring realistic reproduction, physical constraints may be strictly applied, while in content where exaggerated staging is permitted, the criteria for some constraints may be relaxed or specific conditions may be disabled.

[0172] According to some embodiments of the present disclosure, predefined constraints may include predefined physical constraints, which may mean conditions for evaluating whether they are satisfied by performing physical operations on simplified objects that approximate the location or operation of an object within a virtual space. That is, the constraint evaluation in step (S400) may be performed not by retrospectively examining pixel-based results generated by a generative artificial intelligence model, but by determining whether the spatial state and operation state represented by the second data set are physically possible.

[0173] Specifically, referring to FIG. 4, the processor (110) may place a simplified object that approximates the location or operation of an object within a virtual space (S410). Here, the virtual space may refer to a simulation environment or digital space defined to allow computation or simulation to be performed according to physical constraints. The virtual space may be implemented as a two-dimensional or three-dimensional space, and may define a coordinate system (e.g., world coordinate system), gravity direction, unit scale, collision detection rules, physical property parameters (e.g., friction coefficient, mass), boundary conditions (e.g., floor surface, wall surface), etc. Additionally, the virtual space may be created corresponding to a single scene, or may be dynamically updated corresponding to multiple scenes or multiple time intervals. For example, if the content to be created includes multiple time intervals, the virtual space may be configured as a simulation space in which the initial state of objects is set for each time interval and the state is updated over time.

[0174] According to some embodiments of the present disclosure, the virtual space does not require a complete high-resolution 3D scene and may be configured to include only the minimum information necessary for physical consistency evaluation. For example, instead of a complex background mesh, it may include only a plane representing the floor and a simple box shape as an obstacle, and instead of a detailed mesh of a character, it may include a rigging structure approximated by a capsule or a plurality of boxes. Such a minimized virtual space configuration can improve the efficiency of subsequent operations and enable faster detection of physical errors.

[0175] Meanwhile, in the present disclosure, a simplified object may refer to an object in which the location or motion of an object included in the second data set is approximated and represented in a form suitable for physical calculations. Rather than reproducing the shape of the original object (e.g., character, prop, background structure) appearing in the content to be generated exactly as it is, the simplified object may be represented by a simple geometric shape (e.g., box, sphere, capsule, cylinder, or a combination thereof) to facilitate collision detection and physical state updates. For example, a human character may be represented as a set of simplified objects by approximating each part—head, torso, arms, and legs—as a capsule or box; a desk may be represented by approximating the tabletop and legs as a box; and small objects such as a cup may be approximated as a sphere or cylinder. Additionally, the simplified object may be represented as a single object, or a single original object may be broken down into multiple simplified objects.

[0176] A simplified object may include motion information as well as position information. For example, a simplified object may include motion parameters such as position, orientation, and size over a specific time interval, as well as velocity, acceleration, rotational velocity, joint angles (or joint limit ranges), and contact status. Additionally, a simplified object may include physical property parameters for physical calculations. For instance, mass, coefficient of friction, coefficient of restitution, kinematic / dynamic status, and rigid / flexible body status may be assigned to the simplified object, and these parameters may be set based on predefined constraints or object attributes (attribute information included in the first or second data set).

[0177] The processor (110) can place a simplified object in a virtual space using the position and motion information of an object included in the second data set. For example, if the position of object A in the second data set is defined as (x, y, z), the processor (110) can place the simplified object A' by setting its center coordinates to (x, y, z). Additionally, if the orientation of object A is included in the second data set, a rotation matrix or quaternion parameter of the simplified object A' can be set. If the motion of the object is defined by time intervals, the simplified object can be placed so that its position and motion state are updated for each time interval. In this case, the concept of placement can be understood to include not only initial placement at a single point in time, but also the act of setting initial conditions for performing a simulation.

[0178] Next, the processor (110) can determine whether the position or motion of an object satisfies the predefined constraints by performing operations on the simplified object according to predefined physical constraints (S420). Here, the predefined physical constraints may refer to conditions for evaluating whether the physical state of the simplified object does not violate physical laws or rules of physical interaction. The predefined physical constraints may include conditions for evaluating at least one of the possibility of collision between the simplified object and other objects, change in the position of the simplified object due to gravity, restriction on the relative movement of the simplified object due to friction, or continuity of motion of the simplified object due to inertia.

[0179] Conditions for evaluating the likelihood of a collision with other objects may include conditions for determining spatial intersection, clipping, or non-permissible contact between simplified objects. For example, if a simplified object corresponding to a character's hand penetrates a simplified object corresponding to a desk to a depth greater than a certain depth, this may be judged as a violation of the collision condition. In some other embodiments, a threshold distance between objects is defined, and if the distance between simplified objects becomes smaller than that distance, it may be evaluated as a risk of collision or a collision occurring. Such evaluation of collision conditions may be performed as a static judgment at a single frame (single point in time), or it may be performed by predicting the likelihood of a collision occurring at a future point in time through dynamic simulation over a time interval. For example, if the speed and movement path of an object are defined, the processor (110) may predict the position at the next time step and evaluate whether a collision will occur at the predicted position.

[0180] Conditions for evaluating positional changes due to gravity may include conditions for determining whether the simplified object is physically consistent in its state of falling in the direction of gravity or being supported by a support surface. For example, if an object (cup) that should be placed on a table is evaluated to be floating in the air under gravity conditions, this may be determined as a violation of gravity conditions. In some embodiments, if an object continues to penetrate in the direction of gravity or experiences strange positional changes despite being supposed to stop in contact with the floor or support surface, this may be evaluated as a physically inconsistent state. Such gravity condition evaluation may be performed based on the mass of the simplified object, whether it is fixed, gravity acceleration values, etc., and may vary depending on environmental settings, such as setting the gravity value to 0 or a different direction in specific content (e.g., zero-gravity scenes).

[0181] Conditions for evaluating relative movement restrictions based on friction may include conditions for determining whether the relative movement between contacting objects is restricted by friction conditions. For example, if a box placed on a floor continues to slide even without an external force, it may be evaluated as a movement that conflicts with friction conditions. According to some embodiments, when a specific object (e.g., a shoe) moves while in contact with the floor, the degree of sliding should be limited according to the coefficient of friction; if the motion state of the simplified object deviates significantly from this, it may be determined as a violation of friction conditions. Friction condition evaluation may be applied by distinguishing between static friction and kinetic friction, and the coefficient of friction may be set through the material properties of the object or a predefined material-friction mapping table.

[0182] Conditions for evaluating motion continuity based on inertia may include conditions for determining whether a change in velocity or direction of a simplified object is within a physically permissible range. For example, if a person running suddenly drops abnormally to zero in the next time step, or conversely, changes to a very large velocity instantaneously from a stationary state, it may be evaluated as unnatural motion in light of inertia conditions. According to some embodiments, if the angular velocity in rotational motion changes discontinuously or the direction of movement changes abruptly without a physical cause, it may also be judged as a violation of inertia conditions. The evaluation of inertia conditions may be performed using kinematic parameters such as velocity, acceleration, and the rate of change of acceleration over time intervals, and the permissible range of change may vary depending on the type of content. For example, the permissible range of inertia conditions may be set narrowly for realistic content, and wide for exaggerated animation content.

[0183] The processor (110) can perform operations on a simplified object by selectively applying at least one of the physical constraints. For example, if simple static placement verification is required, collision conditions and gravity conditions can be applied first to evaluate static collision and floating, and if the content includes motion, inertia conditions and friction conditions can be additionally applied to evaluate motion continuity over time. Additionally, depending on the characteristics of the scene, they can be applied selectively, such as applying only collision conditions and not gravity conditions in a specific scene. Thus, the selective application of constraints can be implemented by activating or deactivating some of the predefined set of constraints.

[0184] Methods for performing operations according to predefined physical constraints may include various implementations. In one embodiment, the processor (110) may determine whether objects penetrate each other by performing a collision detection algorithm for simplified objects, determine whether the position of an object aligns with a support surface after a certain period of time by performing a position update operation based on gravity, determine whether relative movement in a contact state exceeds an allowable range by applying a friction force model, and determine whether discontinuity of velocity or acceleration exceeds a standard by applying an inertia-based equation of motion. In another embodiment, the processor (110) may call a physics engine or simulation module to run a simulation for a certain period of time, and determine whether collision events, falling events, sliding events, or discontinuous motion events are satisfied by extracting them from the simulation results. However, the present disclosure is not limited to a specific physics engine or a specific numerical integration method.

[0185] Meanwhile, according to some embodiments of the present disclosure, the processor (110) may determine whether the position of an object or the behavior of an object satisfies the predefined constraints by determining whether an evaluation value or event calculated according to predefined physical constraints exists within an allowable range. For example, in the case of a collision evaluation, the processor (110) may determine that it is satisfied if the penetration depth is 0 or less than or equal to a reference depth, and may determine that it is not satisfied if it exceeds the reference depth. In the case of a gravity evaluation, the processor (110) may determine that it is satisfied if, after a certain period of time, the position of the object stops stably on a support surface or exists within an allowable vibration range, and may determine that it is not satisfied if it continues to fall or floats abnormally. In the case of a friction evaluation, the processor (110) may determine that it is satisfied if the relative speed in a contact state does not exceed an allowable range, and may determine that it is not satisfied if it continuously slides without an external force. In the case of inertia evaluation, the processor (110) may determine that the condition is satisfied if the amount of change in velocity or the amount of change in acceleration does not exceed a reference value, and may determine that it is not satisfied if it changes rapidly and discontinuously. Additionally, when multiple constraints are applied, the processor (110) may apply various judgment logics, such as determining that the condition is finally "satisfied" only when all conditions are satisfied, or determining that it is not satisfied immediately if a condition of high importance (e.g., a collision condition) is violated.

[0186] According to some embodiments of the present disclosure, the constraint evaluation process in step (S400) may include a Sim2Real feedback loop. For example, if a physical error such as penetration between objects (e.g., a hand penetrating a desk) is detected, the physics engine may perform inverse kinematics operations to calculate correct coordinate values ​​(constraints) that satisfy the physical constraints and return said coordinate values ​​to the orchestrator. The orchestrator may be controlled to automatically correct the position of an object or the motion of an object included in a second data set based on said coordinate values, and may be configured to re-perform the evaluation of step (S400) on the corrected second data set. Additionally, said coordinate values ​​or the correction results derived from said coordinate values ​​may be reflected in the extraction of analysis data in step (S500) and the conversion of control information in step (S600).

[0187] Referring again to FIG. 2, if the processor (110) determines that the second data set does not satisfy a predefined constraint (S400, No), that is, if it is determined that the location and operation of the object do not satisfy the predefined constraint, it may extract analysis data including the degree of discrepancy with the predefined constraint and the location where the discrepancy occurred (S500). Step (S500) may be defined as a step for quantitatively and spatially describing how the discrepancy occurred, at what point, and to what extent, rather than ending the physical constraint evaluation result with a simple binary judgment (satisfied / unsatisfied).

[0188] According to some embodiments of the present disclosure, the processor (110) can analyze intermediate results produced during a computation or simulation process based on predefined physical constraints to identify the object, time interval, and spatial location where the constraint violation occurred. For example, if an intersection occurs during a collision detection process between simplified objects, the processor (110) can identify which pair of objects the collision occurred between, when the time or time interval in which the collision occurred, and where the spatial coordinates in which the collision occurred are. The information thus identified can be used to specify the location where the discrepancy occurred.

[0189] Additionally, the processor (110) can quantitatively calculate the degree of mismatch. Here, the degree of mismatch may refer to an indicator indicating how severely a predefined constraint has been violated. For example, in the case of a collision constraint, the degree of mismatch may be expressed as at least one of the penetration depth between simplified objects, the size of the overlapping volume, or the length of time the collision is maintained. If the penetration depth is small and occurs instantaneously, the degree of mismatch may be evaluated as low, and if the penetration depth is large or persists for a certain period of time or longer, the degree of mismatch may be evaluated as high.

[0190] In the case of gravity-based constraints, the degree of discrepancy can be expressed as the height at which the object is maintained without a support surface, the time during which no change in position occurs in the direction of gravity, or the difference between the expected trajectory of fall and the actual position. For example, if an object is maintained suspended above a certain height in a virtual space where gravity conditions are applied, that height or the duration of maintenance can be calculated as a value representing the degree of discrepancy.

[0191] In the case of constraints based on friction, the degree of discrepancy can be expressed as the extent to which the allowable relative velocity of movement in contact is exceeded, or the difference between the amount of movement that must be limited by the coefficient of friction and the actual amount of movement.

[0192] In the case of constraints due to inertia, the degree of discrepancy can be calculated as the extent to which the change in velocity or acceleration exceeds the allowable range, or as a numerically expressed value representing the discontinuity of the motion state between time intervals.

[0193] According to some embodiments of the present disclosure, the degree of discrepancy may be expressed as a single numerical value or as a vector composed of multiple parameters. For example, when a collision and a gravity violation occur simultaneously, the collision intrusion depth and the floating height may be calculated as separate values ​​to express the degree of discrepancy multidimensionally. Additionally, the degree of discrepancy may be normalized and expressed as a value between 0 and 1, or expressed as a relative ratio value based on a predefined maximum allowable value.

[0194] Information regarding the location where a mismatch occurred can be expressed as location information in a spatial coordinate system. For example, the location of the mismatch can be expressed as (x, y, z) coordinates within virtual space, or as a specific area in the object's local coordinate system or object surface coordinate system. Additionally, the location of the mismatch can be expressed as a single point, but it can also be expressed as area information (e.g., volume, area, or bounding box) representing the entire area where the mismatch occurred. For example, if a collision between objects occurs across a specific contact area, the entire contact area can be recorded as the location of the mismatch.

[0195] According to some embodiments of the present disclosure, the processor (110) may track the location where a discrepancy occurs according to a time interval and include temporal information, including the time of occurrence, duration, and end time of the discrepancy, together in the analysis data. For example, if a collision occurs only in a specific interval while a specific object is moving, the time when the collision started and the time when it ended may be recorded, and this may be used to determine in a subsequent step which interval the discrepancy occurs intensively.

[0196] The processor (110) can generate analysis data by combining the degree of mismatch and the location of the mismatch calculated as above. The analysis data may consist of structured data including the type of constraint violation, the object where the violation occurred, the location where the violation occurred, the severity of the violation, and the time interval where the violation occurred. For example, the analysis data may include information in the form of "A collision with an intrusion depth of 0.15 m occurred between object A and object B during the time interval t2 to t3." Or it may include information in the form of "Object C was stationary at a height of 0.3 m for 1.2 seconds under gravity conditions."

[0197] According to some embodiments of the present disclosure, analysis data may be configured in a specific data structure or format so that it is easy to convert into control information interpretable by a generative artificial intelligence model in a subsequent step. For example, the analysis data may be configured in a mapped form of numerical data indicating the degree of discrepancy and coordinate data indicating the location of the discrepancy, which can be used as an input for conversion into map data, mask data, or weight data in a subsequent step. Additionally, the analysis data may be configured as information regarding a single frame or a single point in time, or as time-series data including information regarding multiple time intervals.

[0198] Meanwhile, if analysis data is extracted in step (S500), the processor (110) can convert the analysis data into control information that can be interpreted by a generative artificial intelligence model (S600). Step (S600) can be defined as a step that does not merely detect discrepancies with predefined constraints, but converts the detected discrepancies so that they can be directly reflected in the generation process of the generative artificial intelligence model.

[0199] According to some embodiments of the present disclosure, the control information may be generated based on a coordinate value (Constraint) calculated by the physics engine by inverse kinematics operation in step (S400), or a combination of said coordinate value and constraint violation information.

[0200] In the present disclosure, control information may refer to data expressed in a form that allows a generative artificial intelligence model to interpret and reflect as a condition during the generation process information regarding the degree of discrepancy with predefined constraints and the location where the discrepancy occurred.

[0201] Control information can be configured to induce selection tendencies in the generation process or suppress the possibility of specific outcomes without directly modifying the internal structure or learning parameters of the generative AI model. In other words, control information functions as an input condition or conditional control signal for the generative AI model and can be utilized as an intermediate control means capable of influencing the entire generation process.

[0202] According to some embodiments of the present disclosure, control information may include at least one of map data indicating the degree of discrepancy with a predefined constraint and the location where the discrepancy occurred, mask data generated based on the map data and specifying a part of the content to be generated, and weight data calculated based on the degree of discrepancy with the predefined constraint. This control information may be used alone or a combination of multiple control information, and may be selectively applied depending on the type of content to be generated, the type of generative artificial intelligence model, or the strength of automatic correction.

[0203] In the present disclosure, map data may refer to data that spatially represents the location where a discrepancy with a predefined constraint occurs and the degree of such discrepancy. Map data may be expressed in the form of a two-dimensional map or a three-dimensional map, and discrepancy information may be represented in a distribution form on a coordinate system corresponding to the spatial structure of the content to be generated. For example, map data may include an error heatmap having a value representing the degree of discrepancy for each location on a coordinate system corresponding to an image or video frame. In this case, the degree of discrepancy may be expressed as a color, brightness, saturation, or numerical value, and areas with more severe discrepancies may be indicated with higher values ​​or stronger colors.

[0204] According to some embodiments of the present disclosure, map data may be configured in the form of a texture map that visualizes the degree of discrepancy with a predefined constraint, i.e., the penetration depth, as a color gradient.

[0205] For example, areas with a calculated low intrusion depth are represented by low saturation or colors at one extreme, and areas with a calculated high intrusion depth are represented by high saturation or colors at the other extreme, thereby allowing the location and severity of the discrepancy to be spatially and intuitively expressed.

[0206] Such color gradient-based texture maps can be mapped onto a 2D image coordinate system or a 3D surface coordinate system, provided as conditional input to a generative artificial intelligence model, or used as reference data for generating mask data in a subsequent step.

[0207] The processor (110) can generate map data using information regarding the location of the discrepancy and the degree of discrepancy included in the analysis data extracted in step (S500). Here, the degree of discrepancy may refer to the penetration depth. For example, if a collision between simplified objects occurs in a specific coordinate area, a value proportional to the degree of discrepancy may be assigned to the location of the map data corresponding to that coordinate area. A higher value may be assigned as the collision penetration depth increases, and a lower value may be assigned as the penetration depth decreases. Additionally, in the case of a gravity condition violation, the height or duration of the object being held in the air may be converted into the degree of discrepancy and reflected in the map data corresponding to that location. In this way, the map data may be a spatial representation that simultaneously includes the location and degree of discrepancy.

[0208] In the present disclosure, mask data may refer to data generated based on map data and specifying a part of the content to be generated. The mask data may consist of binary or multi-valued data for distinguishing between areas requiring automatic correction and areas that do not, among the entire area of ​​the content to be generated. For example, only areas where the degree of inconsistency exceeds a predefined threshold in the map data may be selected and converted into mask data; in this case, the mask data may be configured in the form of a binary mask that indicates "areas to be corrected" as 1 and "areas not to be corrected" as 0.

[0209] The processor (110) can generate mask data by considering the value distribution of the map data, the type of mismatch, or the duration of the mismatch. For example, areas where collisions occur temporarily may be excluded from the mask, and only areas where collisions have occurred for a certain period of time or longer may be designated as the mask. Additionally, a multi-value mask may be generated by partially including areas with a low degree of mismatch in the mask for soft correction, and definitely including areas with a high degree of mismatch in the mask for strong correction. The mask data generated in this way can be combined with the in-painting, region regeneration, or local conditional generation functions of a generative AI model to induce selective correction of only the problem area.

[0210] In the present disclosure, weight data refers to numerical data calculated based on a numerical value that quantifies the degree of discrepancy with a predefined constraint, i.e., the penetration depth, and may refer to data applied to negative prompt tags or prompt elements included in the input of a generative AI model to control the degree of selection or suppression of a specific generated result. Weight data can serve to numerically express how strongly a generative AI model should avoid a specific feature or pattern during the generation process.

[0211] According to some embodiments of the present disclosure, the penetration depth may be quantitatively calculated from a collision evaluation between simplified objects or from the result of calculations based on physical constraints.

[0212] For example, the intrusion depth can be calculated as at least one of the distance at which the surfaces or boundaries of two simplified objects overlap each other, the ratio of the overlapping volume, or the cumulative value of the overlapping state maintained over a certain time interval.

[0213] In addition, the above-mentioned intrusion depth may be expressed in physical distance units (e.g., meters) or converted into a dimensionless value normalized to a range of 0.0 to 1.0 so as to be suitable for application as input to a generative artificial intelligence model.

[0214] According to some embodiments of the present disclosure, weight data may be defined as a value obtained by back-calculating the penetration depth calculated in the physical constraint evaluation step into a natural language prompt weight that can be interpreted by a generative artificial intelligence model. According to some embodiments of the present disclosure, weight data may be calculated by the following formula based on the degree of discrepancy with a predefined constraint, i.e., the penetration depth.

[0215]

[0216] Here, Weight represents a weight to be applied to a negative prompt tag or prompt element included in the input of a generative artificial intelligence model, and can be set to a value in the range of, for example, 1.0 to 2.0.

[0217] Depth represents the depth of intrusion between collision objects and can be expressed in units of physical distance (e.g., meters) or as a value normalized to a range of 0.0 to 1.0.

[0218] a (alpha) represents a depth-weighted conversion coefficient for converting the intrusion depth into a weight value, and can be set to a value such as 10.0, for example.

[0219] B (beta) represents a weight offset that is applied by default even when the intrusion depth is close to 0, and can be set to a value such as 0.5, for example.

[0220] According to the above formula, as the intrusion depth increases, the weight data increases linearly, which can have the effect of inducing the generative AI model to reflect the corresponding negative prompt tag or prompt element more strongly during the generation process. For example, when the intrusion depth is relatively small, a low weight is applied so that the degrees of freedom of the generation result can be maintained at a certain level, and when the intrusion depth is large, a high weight is applied so that patterns causing penetration between objects, abnormal placement, or physical errors can be strongly suppressed during the generation process.

[0221] For example, if the depth of intrusion between collision objects is calculated to be 0.1m as a result of the predefined constraint evaluation, the weight data according to the above formula It can be calculated as such. The weight data calculated in this way can be inserted into a prompt in the form of (intersecting objects: 1.5) as in some embodiments of the present disclosure. Accordingly, the generative AI model can be induced to strongly suppress features or patterns related to penetration between objects during the generation process. Consequently, the processor (110) can convert the penetration depth value included in the analysis data in step (S600) into a negative prompt weight through numerical calculation and automatically assign the negative prompt weight to the negative prompt tag of the generative AI model.

[0222] The present disclosure is not limited to the above formula, and the values ​​of a (alpha) and B (beta) may be predefined or dynamically adjusted depending on the type of generative artificial intelligence model, the type of content to be generated, the importance of the scene, or the strength of automatic correction.

[0223] Meanwhile, according to some embodiments of the present disclosure, the linear function according to the above formula is not limited, and a non-linear function for calculating weight data in proportion to the penetration depth may also be applied.

[0224] For example, weight data can be calculated to increase non-linearly with respect to the penetration depth by an exponential function, a logarithmic function, or a piecewise function.

[0225] According to some embodiments of the present disclosure, when the calculated weight data is greater than or equal to a reference value of 1.0, the suppression strength for the corresponding negative prompt tag or prompt element may be controlled to be strengthened during the generation process of the generative artificial intelligence model.

[0226] That is, when the weight data is less than 1.0, the condition is reflected gradually, and when the weight data is 1.0 or higher, the generation path can be adjusted so that features or patterns causing physical errors are preferentially excluded during the generation process.

[0227] Additionally, if the intrusion depth is calculated at multiple locations or multiple time intervals, a weight corresponding to each intrusion depth may be calculated individually and applied to different negative prompt tags or prompt elements.

[0228] In other words, weight data functions as a control signal that automatically corrects the generation results by numerically reflecting the severity of physical constraint violations, enabling gradual and precise generation control rather than simple binary blocking.

[0229] Here, a negative prompt tag refers to a conditional element included in the input prompt of a generative AI model that induces the model not to generate specific attributes, forms, behaviors, or patterns. For example, in the case of an image generation model, negative prompt tags such as "intersecting objects," "bad anatomy," and "floating object" may be used, and these tags can induce the model not to select results containing those characteristics. Meanwhile, a prompt element refers to an individual word, phrase, sentence, or conditional item that constitutes the input prompt of a generative AI model, and can include both positive and negative elements.

[0230] The processor (110) can calculate weight data using the degree of discrepancy included in the analysis data. For example, if the degree of discrepancy is expressed as a numerical value, a weight proportional to that numerical value may be calculated, and this weight may be applied by multiplying or adding to a negative prompt tag or a specific prompt element. According to some embodiments of the present disclosure, the greater the degree of discrepancy, the greater the weight applied to the negative prompt tag, so that the tag may act more strongly during the generation process. Conversely, if the degree of discrepancy is small, the weight may be set low, so that the influence of the tag may be relatively weakened.

[0231] The application of weight data to negative prompt tags or prompt elements can mean that the probability or priority of selecting or suppressing specific features is adjusted during the generation process of a generative AI model. For example, if a generative AI model internally evaluates multiple candidate results to select a final outcome, applying a high weight to a negative prompt tag may lower the score or reduce the selection probability of candidate results containing features associated with that tag. Conversely, if the weight is low, the influence of the tag weakens, allowing the model to perform more free variations.

[0232] According to some embodiments of the present disclosure, weight data may be applied collectively to the entire creation process of the content to be created, or it may be applied selectively only to specific time intervals or specific scenes. For example, if a discrepancy occurs only in a specific frame interval, local automatic correction may be achieved by applying weight data only to the prompt element or negative prompt tag corresponding to that interval. Additionally, weight data may be applied together with mask data so that the influence of the negative prompt tag is strengthened only in the area designated by the mask.

[0233] As described above, the map data, mask data, and weight data included in the control information act complementarily to precisely control the extent to which specific regions, features, or behaviors are selected or suppressed during the generation process of the generative AI model. This prevents the repeated generation of results that violate predefined constraints and enables the automatic generation of content with enhanced physical consistency and behavioral realism.

[0234] Meanwhile, if the processor (110) converts the analysis data into control information in step (S600), it can generate final content by providing the control information as a conditional input to a generative artificial intelligence model (S700). Step (S700) can be defined as a step that not only discards the content to be generated that includes elements violating predefined constraints, but also automatically generates corrected content based on the verification results.

[0235] In this disclosure, a conditional input may refer to an input that, during the process of generating final content, does not merely perform probabilistic sampling based solely on input data (e.g., text prompts), but rather incorporates one or more additionally provided constraints or guidance information to restrict or induce the spatial composition, selection tendencies, or stylistic identity of the generated results. That is, the conditional input may function as an input that influences the distribution of candidate results, selection probabilities, or the direction of change of latent representations during the generation process of the generative AI model, and in this disclosure, control information may be provided as a conditional input. In other words, since the control information is generated based on analysis data including the location and degree of violation of predefined constraints, providing it as a conditional input may induce the generative AI model to avoid or modify areas or patterns that cause physical errors.

[0236] According to some embodiments of the present disclosure, a processor (110) can control the generation process so that the generation result satisfies predefined constraints by maintaining the semantic direction of the content that the generative AI model is to generate based on input data, while also providing conditional inputs such as control information, guide data, reference images, and appearance characteristic data. That is, while the input data provides semantic and narrative goals regarding "what to generate," the conditional inputs can function to provide structural, physical, and identity constraints regarding "how to generate." In this case, the conditional inputs can provide substantial control over the generation process without changing the internal structure of the generative AI model.

[0237] Specifically, referring to FIG. 5, the processor (110) can generate guide data by extracting the spatial location of an object, the relative placement of an object, and the operational state of an object from a second data set (S710). Here, guide data may refer to data constructed by extracting or converting information that defines the spatial and structural framework that a generative AI model must follow during the generation process, among the location and operational information of an object included in the second data set. The guide data does not directly define the pixel value of the final content itself, but can function as a guideline to limit or guide the generation space of the generative AI model so that the placement of objects, scene composition, development of actions, or temporal flow do not violate constraints.

[0238] According to some embodiments of the present disclosure, a processor (110) may generate guide data using coordinate information of an object included in a second data set, relative distance relationships between objects, relative positional relationships such as up, down, left, and right, and operation state information for each time interval. For example, if the content to be generated corresponds to a single frame or a single scene, the guide data may be composed of two-dimensional or three-dimensional layout information representing the approximate location and placement relationship of each object. On the other hand, if the content to be generated corresponds to multiple frames of video or multiple time intervals, the guide data may be generated in the form of a sequence representing changes in the position and operation state of an object for each time interval. In this case, the guide data may include position and operation states sampled at regular time intervals, or may include state information sampled more densely in intervals where contact or collision is likely to occur.

[0239] Additionally, guide data can be generated by converting it into a representation that can be utilized by a generative AI model. For example, guide data can be generated in the form of a structure map representing the arrangement of objects, a depth map representing depth information, a pose map representing the posture or movement of objects, or a combination thereof. Such map data can be utilized as structural conditions to guide the generative AI model to follow the spatial composition or the skeletal structure of the movement during the generation process. However, the present disclosure is not limited to a specific format of guide data, and guide data may be configured in various formats depending on the type of generative AI model.

[0240] Meanwhile, when guide data is generated (S710), the processor (110) can generate final content based on input data by providing control information along with a pre-registered reference image or appearance characteristic data as a conditional input to a generative AI model (S720). Here, a reference image may refer to image data used to maintain the appearance, style, background atmosphere, color tone, or texture of a specific object or character appearing in the content to be generated. For example, a face image, clothing image, background image, or reference image representing a specific style of a specific character may be registered as a reference image, and this can be used as standard information to enable the generative AI model to consistently maintain the appearance or style during the generation process.

[0241] In addition, appearance characteristic data may refer to feature vectors, embeddings, keypoints, or sets of parameters representing the appearance extracted from a reference image, which are used to enable a generative AI model to reliably reproduce the appearance. For example, appearance characteristic data may include feature points of a character's face, features of clothing patterns, color palettes, or embedding vectors representing style characteristics.

[0242] According to some embodiments of the present disclosure, the reference image or external characteristic data may be stored in a storage unit (120) or managed in a state of being stored in an external server or external storage accessible through a communication unit (130).

[0243] Reference images or appearance characteristic data may be registered directly by the user, or they may be data pre-registered by a system administrator or content creator. For example, the user may upload a specific character's face image, clothing image, or background style image as a reference image and save it to the storage unit (120). Additionally, corporate users or content service providers may pre-register brand characters, fixed worldview assets, or repeatedly used background and prop images as appearance characteristic data.

[0244] Additionally, according to some embodiments of the present disclosure, the appearance characteristic data does not necessarily need to be stored in the form of an original image, but may be stored in a summarized form such as a feature vector, embedding, keypoint information, color distribution information, or style parameters extracted from a reference image. In this case, the processor (110) may selectively load and use only the appearance characteristic data in a form suitable for providing as a conditional input to a generative artificial intelligence model.

[0245] Reference images or appearance characteristic data can be managed by being classified by user account unit, project unit, or content type unit, and can be selectively applied by the processor (110) depending on the type of content to be created or the scene composition. For example, if the same character appears in multiple scenes, the same appearance characteristic data can be repeatedly applied to maintain the character's identity.

[0246] In step (S720), the role of conditional input is not merely to provide additional information, but to limit probabilistic fluctuations that may occur during the generation process of the generative AI model and to guide the generation path so that areas or patterns where physical constraint violations occurred are not reproduced. For example, map data and mask data of control information can act as spatial conditions that allow the model to recognize areas where discrepancies occurred and cause those areas to be regenerated or modified. Additionally, weight data can be applied to negative prompt tags or prompt elements to reduce the probability that features causing physical errors (e.g., penetration, floating, discontinuous motion) are selected during the generation process. Guide data can function as structural conditions that define the structural framework of the object's placement and motion, ensuring that location, placement, and motion states satisfying the constraints are reflected in the generation result. Reference images or appearance characteristic data can function as identity conditions that maintain the character's identity, appearance, and style. Ultimately, conditional inputs generate a result in which "structure, physicality, and identity" are corrected while maintaining "meaning (input data)" by simultaneously providing multiple conditional signals that perform different roles.

[0247] According to some embodiments of the present disclosure, a generative artificial intelligence model may be configured to reflect conditional inputs at different stages of the generation process. For example, the overall composition and layout may be determined by guide data at the beginning of generation, regeneration may be performed in areas where mask data is applied during the generation process, and negative prompt tags with weighted data may act to suppress the occurrence of specific patterns throughout the generation process. Additionally, the consistency of the appearance of the generated result may be maintained by a reference image or appearance characteristic data. In this way, conditional inputs may operate in a multi-layered manner throughout the generation process.

[0248] When final content is generated in this manner, the generative AI model can produce a result in which elements that violated predefined constraints are suppressed or modified, while maintaining semantic and narrative goals based on input data. That is, the final content has a reduced probability of errors such as penetration, floating, unrealistic movement, and motion discontinuity that violate physical constraints, while simultaneously maintaining object placement, motion development, and identity consistency through the input of guide data and appearance conditions. Therefore, step (S700) according to some embodiments of the present disclosure can provide the effect of achieving automatic correction without relying on post-editing or random retries by controlling the generation process itself by reducing the verification result to a conditional input, rather than simply performing regeneration after verification.

[0249] In short, the processor (110) is configured to provide control information as a conditional input to a generative artificial intelligence model, and simultaneously provides guide data generated from a second data set and pre-registered reference images or external characteristic data together as conditional inputs, thereby generating final content based on input data, guiding the generation path so that violations of pre-defined constraints do not recur, and achieving the effect of improving physical consistency and spatiotemporal consistency.

[0250] Meanwhile, according to some embodiments of the present disclosure, guide data may be generated in correspondence with multiple time intervals of the content to be generated. For example, if the content to be generated is video content comprising multiple frames, the guide data may be generated in frame units, scene units, or in predefined time interval units. In this case, the guide data corresponding to each time interval may be configured to reflect the location, placement, and operation state of an object within that interval.

[0251] In addition, control information may also be generated corresponding to multiple time intervals, or generated in a form that applies commonly to said multiple time intervals. For example, if a violation of physical constraints occurs only in a specific time interval, only the control information corresponding to that time interval may be generated and selectively applied. On the other hand, if the same constraint is repeatedly violated throughout the content to be generated, the control information may be generated to apply commonly to the entire time interval.

[0252] According to some embodiments of the present disclosure, if collisions between objects occur only in specific frame intervals, map data and mask data corresponding to said frame intervals are generated, and weight data may also be applied only to said intervals. In some other embodiments, if discontinuity in character motion occurs throughout the content, weight data may be applied commonly across the entire time interval to induce a generative AI model to reinforce the overall continuity of motion.

[0253] In this way, as guide data and control information are applied separately by time intervals or universally, the generative AI model may be subject to strict constraints in specific sections of the generation process and relatively relaxed constraints in others. Through this, it is possible to selectively modify only the areas requiring automatic correction while maintaining the overall flow and naturalness of the content.

[0254] Meanwhile, referring again to FIG. 2, if the processor (110) determines in step (S400) that the second data set satisfies predefined constraints (S400, Yes), that is, if it determines that the location of the object and the operation of the object satisfy predefined constraints, it can generate guide data by extracting the spatial location of the object, the relative placement of the object, and the operation state of the object from the second data set, and provide the guide data and the pre-registered reference image or external characteristic data as conditional inputs to a generative artificial intelligence model to generate the final content based on the input data (S800).

[0255] Guide data may include data that directly reflects the spatial location of objects included in the second dataset, the relative placement relationships between objects, and the operational state of objects, or data converted into a form that is easy for a generative AI model to interpret. Meanwhile, reference images or appearance characteristic data may be used to maintain the appearance identity of objects or characters. As the descriptions of guide data, reference images, and appearance characteristic data have been described above, a detailed explanation is omitted.

[0256] Consequently, when the processor (110) determines that the second data set satisfies predefined constraints, it can generate final content based on input data by providing guide data and reference images or appearance characteristic data as conditional inputs without additional automatic correction or regeneration control.

[0257] The generative AI content generation control method described above can provide the effect of simultaneously improving generation quality and computational efficiency by branching the generation path according to whether predefined constraints are satisfied, thereby omitting unnecessary correction steps and enabling efficient generation when the constraints are satisfied, and generating a result with physical consistency through an automatic correction route when the constraints are not satisfied.

[0258] According to at least one of the embodiments of the present invention described above, in the process of generating content by a generative artificial intelligence model, semantic verification and physical verification according to predefined constraints are performed stepwise, and automatically corrected content is generated based on the verification results, thereby providing the advantage of effectively preventing the generation of content containing errors such as physical hallucinations, motion discontinuities, object penetration, or narrative inconsistencies. Furthermore, according to some embodiments of the present disclosure, if the verification is passed, unnecessary correction steps can be omitted and the final content can be generated immediately, thereby providing the effect of reducing waste of computational resources and improving overall generation efficiency while maintaining generation quality.

[0259] In the present disclosure, the device (100) is not limited to the configuration and method of the several embodiments described above; rather, all or part of each embodiment may be selectively combined to allow for various modifications to be made. For example, some of the semantic verification step, physical verification step, analysis data extraction step, control information conversion step, and final content generation step may be selectively omitted or performed in combination with other steps, and such modifications should also be understood as being included within the technical scope of the present invention.

[0260] The various embodiments described in this disclosure may be implemented, for example, in a recording medium readable by a computer or similar device using software, hardware, or a combination thereof. According to hardware implementations, some embodiments described herein may be implemented using at least one of an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing unit (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a processor, a controller, a microcontroller, a microprocessor, or other electrical unit for performing functions. According to some embodiments, the configurations and methods described in this disclosure may be implemented to be executed by at least one processor.

[0261] According to a software implementation, some embodiments such as the procedures and functions described in this disclosure may be implemented by a plurality of software modules, each of which may be configured to perform one or more functions, tasks, or operations described in this disclosure. The software modules may be implemented by software code written in a suitable programming language, said software code may be stored in a storage unit (120) and executed by at least one processor (110). That is, by storing at least one program instruction in the storage unit (120) and executing said program instruction by at least one processor (110), the verification-based generation control and automatic correction method for generative artificial intelligence content according to the present invention may be performed.

[0262] Additionally, a generative artificial intelligence content generation control method according to some embodiments of the present disclosure may be implemented in the form of code executable by at least one processor on a recording medium readable by at least one processor equipped in the device (100). The recording medium includes all types of recording devices that store data that can be read by at least one processor, and may include, for example, ROM (Read Only Memory), RAM (Random Access Memory), CD-ROM, magnetic tape, magnetic disk, optical disk, or other non-transitory computer-readable recording media.

[0263] Meanwhile, although the present disclosure has been described with reference to the attached drawings, this is merely an example to aid in understanding the invention, and the invention is not limited to specific embodiments. Those skilled in the art to which the invention pertains may make various modifications and changes based on the technical concept and claims of the present invention, and such modifications and changes should also be interpreted as being included within the scope of the rights of the present invention. Furthermore, such modified embodiments should not be understood separately from the technical concept of the present invention, but should be understood integrally within the scope of the overall technical concept of the present invention.

Claims

Claim 1 A method for controlling content generation using a generative artificial intelligence model by a processor of a device, wherein the method comprises: a step of, when input data for content generation is received from a user, processing the input data to generate a first data set including at least one of semantic information, structural characteristic information, or narrative context information regarding content to be generated; a step of determining whether the first data set satisfies the predefined generation conditions by comparing the first data set and the content to be generated with predefined generation conditions in a vector space; a step of, when it is determined that the first data set satisfies the predefined generation conditions, generating a second data set including information on at least one of the location of an object included in the content to be generated or the action of the object based on the first data set; and a step of determining whether the location of the object and the action of the object satisfy the predefined constraints by performing an evaluation according to predefined constraints on the second data set. If it is determined that the location and operation of the object do not satisfy the predefined constraints, the method comprises: a step of extracting analysis data including the degree of discrepancy with the predefined constraints and the location where the discrepancy occurred; a step of converting the analysis data into control information in a form interpretable by the generative AI model; and a step of generating final content by providing the control information as a conditional input to the generative AI model; wherein the step of converting the analysis data into control information interpretable by the generative AI model comprises: a step of converting a penetration depth value included in the analysis data into a negative prompt weight through numerical calculation; and a step of automatically assigning the negative prompt weight to the negative prompt tag of the generative AI model.A method comprising: a step of determining whether the location of an object and the operation of an object satisfy predefined constraints, wherein the step of determining whether the location of an object and the operation of an object satisfy predefined constraints comprises: a step of placing a simplified object that approximates the location of an object or the operation of an object in a virtual space; and a step of determining whether the location of an object or the operation of an object satisfies the predefined constraints by performing an operation according to predefined physical constraints on the simplified object. Claim 2 A method according to claim 1, wherein the control information comprises at least one of map data including an error heatmap indicating the degree of discrepancy with the predefined constraint and the location where the discrepancy occurred, mask data generated based on the map data including the error heatmap and specifying a part area of ​​the content to be generated, and weight data calculated based on a numerical value quantifying the degree of discrepancy with the predefined constraint. Claim 3 A method according to paragraph 2, wherein the weight data is applied to a negative prompt tag or prompt element included in the input of the generative artificial intelligence model to control the degree to which the negative prompt tag or the prompt element is selected or suppressed during the generation process of the target content. Claim 4 A method according to claim 1, wherein the step of determining whether the predefined generation condition is satisfied comprises: a step of vectorizing the first data set and the predefined generation condition to calculate a similarity in the vector space; and a step of determining whether the predefined generation condition is satisfied based on whether the similarity satisfies a reference value. Claim 5 delete Claim 6 A method according to claim 1, wherein the prior-defined physical constraint includes a condition for evaluating at least one of the possibility of collision between the simplified object and other objects, a change in the position of the simplified object due to gravity, a restriction on the relative movement of the simplified object due to friction, or the continuity of motion of the simplified object due to inertia. Claim 7 The method according to claim 1, wherein the step of generating final content by providing the control information as a conditional input to the generative artificial intelligence model comprises: a step of generating guide data by extracting the spatial location of the object, the relative arrangement of the object, and the operational state of the object from the second data set; and a step of generating the final content based on the input data by providing the control information as a conditional input to the generative artificial intelligence model together with the guide data and a previously registered reference image or appearance characteristic data. Claim 8 A method according to claim 7, wherein the guide data is generated corresponding to a plurality of time intervals, and the control information is generated for each of the plurality of time intervals or is generated in a form that is commonly applied to the plurality of time intervals. Claim 9 The method according to claim 1, further comprising the step of controlling so that subsequent steps for generating the final content are not performed when it is determined that the first data set does not satisfy the predefined generation conditions. Claim 10 A method according to claim 1, further comprising: a step of generating guide data by extracting the spatial location of the object, the relative placement of the object, and the operational state of the object from the second data set when it is determined that the location and operation of the object satisfy the predefined constraints; and a step of generating the final content based on the input data by providing the guide data and a pre-registered reference image or appearance characteristic data as conditional inputs to the generative artificial intelligence model.

Citation Information

Patent Citations

  • Apparatus and method for training artificial intelligence based model for detecting object in driver monitoring systems

    KR1020250074782A

  • Outpainting service providing server, system, method and program that creates images with changed viewpoints using generative ai

    KR1020250113874A

  • Method and apparatus for providing optimized advertising using generative artificial intelligence model

    KR1020250141001A

  • Servers, systems, methods, and programs that provide custom model creation services using generative artificial intelligence

    KR102713202B1

  • Method for generating three-dimensional modeling data of traditional timber structure using ai-based image recognition and computer-readable recording medium having a program stored thereon for implementing the method

    KR102888167B1