Modifying digital scripts using generative adversarial networks

The GAN system addresses inefficiencies in digital script modification by automating text and image sequence changes based on user interactions, improving the interactivity of digital storytelling.

JP7841830B2Active Publication Date: 2026-04-07INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-08-25
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing methods for modifying digital scripts are inefficient and lack automation in dynamically changing the text content based on user interactions and image sequences.

Method used

A generative adversarial network (GAN) system that enables natural language processing (NLP) to identify contextual dimensions, generate and modify image sequences, and dynamically change the text content of digital stories based on user interactions.

Benefits of technology

Automates the modification of digital scripts by allowing for dynamic changes in text content in response to user interactions and image modifications, enhancing the flexibility and interactivity of digital storytelling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007841830000001
    Figure 0007841830000001
  • Figure 0007841830000002
    Figure 0007841830000002
  • Figure 0007841830000003
    Figure 0007841830000003
Patent Text Reader

Abstract

A system, method, and computer program product for implementing modification of a digital script are provided. The method includes generating an image sequence associated with textual content of a digital story. A plurality of contextual dimensions within the textual content are identified and a group of dimensions is selected. The image sequence combined with the group of dimensions is scaled and the image sequence is modified based on the detected interaction with the group of dimensions. During presentation of the digital story, a scriptwriter is enabled to extract dimensions from the group of dimensions and modify the dimensions. The image sequence is modified and a hardware interface device is enabled to interact with the various image sequences and modify the plurality of contextual dimensions. The textual content of the digital story is dynamically modified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to a method for modifying a digital script, and more particularly to a software technique for generating and modifying an image sequence related to the text content of a digital story and dynamically changing the related digital text content, and to a method and related system for improving the same.

Summary of the Invention

[0002] A first aspect of the present invention is a generative adversarial network (GAN) hardware device including a processor coupled to a computer-readable memory unit, the memory unit including instructions that, when executed by the processor, implement a method for modifying a digital script that enables natural language processing (NLP), the method comprising: generating, by the processor, an image sequence related to the text content of a digital story; identifying, by the processor via executing NLP code, a plurality of context dimensions within the text content; selecting, by the processor in response to user input, a group of dimensions of the plurality of context dimensions; expanding or shrinking, by the processor, the image sequence in combination with the group of dimensions; changing, by the processor, the image sequence based on a detected interaction with the group of dimensions; extracting, by the processor, dimensions from the group of dimensions during presentation of the digital story and the image sequence; activating, by the processor, a scriptwriter related to the text content of the digital story to modify the dimensions; modifying, by the processor, the image sequence based on the dimension modification resulting in response to the activation; interacting, by the processor, with various image sequences of the image sequence and activating a hardware interface device to change a plurality of context dimensions; and dynamically changing, by the processor in response to the activation, the text content of the digital story. A generative adversarial network (GAN) hardware device is provided.

[0003] A second aspect of the present invention provides a digital script modification method that enables natural language processing (NLP), comprising: generating an image sequence related to the text content of a digital story by a processor in a generative adversarial network (GAN) hardware device; identifying multiple contextual dimensions within the text content by the processor executing NLP code; selecting a group of dimensions of the multiple contextual dimensions by the processor in response to user input; scaling the image sequence in combination with the group of dimensions by the processor; modifying the image sequence based on detected interactions with the group of dimensions by the processor; extracting dimensions from the group of dimensions during the presentation of the digital story and the image sequence by the processor; activating a script writer related to the text content of the digital story to modify the dimensions by the processor; modifying the image sequence based on the dimension modifications that occur in response to the activation by the processor; activating a hardware interface device to interact with various image sequences of the image sequence and modify multiple contextual dimensions by the processor; and dynamically changing the text content of the digital story in response to the activation by the processor.

[0004] A third aspect of the present invention comprises a computer-readable hardware storage device for storing computer-readable program code, the computer-readable program code including an algorithm that implements a digital script modification method that enables natural language processing (NLP) when executed by a server processor, the method comprising: generating an image sequence related to the text content of a digital story by the processor; identifying multiple contextual dimensions within the text content by the processor through the execution of NLP code by the processor; selecting a group of dimensions of the multiple contextual dimensions by the processor in response to user input; and scaling the image sequence in combination with the group of dimensions by the processor The present invention provides a computer program product that includes modifying an image sequence based on detected interactions with a group of dimensions, extracting dimensions from a group of dimensions during the presentation of a digital story and an image sequence, enabling a script writer associated with the text content of the digital story to modify dimensions, modifying the image sequence based on the dimension modifications that occur in response to the activation, enabling a hardware interface device to interact with various image sequences of the image sequence and modify multiple contextual dimensions, and dynamically changing the text content of the digital story in response to the activation.

[0005] This invention advantageously provides a simple method and related system that can automate the modification of digital scripts. [Brief explanation of the drawing]

[0006] [Figure 1] This invention presents a system for improving software techniques related to generating and modifying image sequences associated with text content of a digital story, and dynamically changing the associated digital text content, according to embodiments of the present invention. [Figure 2] This document details an algorithm that illustrates the process flow enabled by the system in Figure 1, for improving software techniques related to generating and modifying image sequences associated with text content of a digital story and dynamically changing the associated digital text content, according to embodiments of the present invention. [Figure 3] This is an internal structure diagram of the software / hardware according to an embodiment of the present invention. [Figure 4] This invention presents a system, comprising a GAN module and an NLP module, for modifying the digital script of digital story content, according to an embodiment of the present invention. [Figure 5A] This document illustrates a process for modifying a digital script and generating a corresponding image sequence, according to an embodiment of the present invention. [Figure 5B] This document illustrates a process for modifying a digital script and generating a corresponding image sequence, according to an embodiment of the present invention. [Figure 5C] This document illustrates a process for modifying a digital script and generating a corresponding image sequence, according to an embodiment of the present invention. [Figure 5D] This document illustrates a process for modifying a digital script and generating a corresponding image sequence, according to an embodiment of the present invention. [Figure 6] This is a detailed diagram of the text-to-image GAN network component in Figure 5, according to an embodiment of the present invention. [Figure 7] This is a detailed diagram of the image-to-text GAN network component according to an embodiment of the present invention, as shown in Figure 5. [Figure 8] To improve software techniques related to generating and modifying image sequences associated with the text content of a digital story and dynamically changing the associated digital text content, according to embodiments of the present invention, a computer system used by the system in Figure 1 is shown. [Figure 9] This document illustrates a cloud computing environment according to an embodiment of the present invention. [Figure 10] This describes a set of functional abstraction layers provided by a cloud computing environment according to embodiments of the present invention. [Modes for carrying out the invention]

[0007] Figure 1 shows a system 100 for improving software techniques related to generating and modifying image sequences associated with text content of a digital story and dynamically changing the associated digital text content, according to an embodiment of the present invention. A typical application script writing system may require a script writer entity to visualize a video presentation that requires text content analysis for image sequence creation. Similarly, during the aforementioned process, the script writer entity may generate requests to expand the generated image sequence with respect to various dimensions. Furthermore, the requests may include specifications for displaying a summary of the associated content, along with commands to modify the associated images, with respect to updating the text content related to the image modification. Thus, the system is configured to enable the script writer entity to expand the generated image content within various contextual dimensions in order to perform image modification. Similarly, system 100 enables automatic updating of text content with respect to image modification.

[0008] System 100 includes a system that enables natural language processing (NLP) to analyze digital text story content, identify various relevant contextual dimensions, and automatically generate image sequences based on the text story content via the execution of a generative adversarial network (GAN). Similarly, System 100 is configured to enable a scriptwriter entity to scale, modify, or bother with the generated images in order to dynamically update the digital text story content. System 100 enables the following functions:

[0009] System 100 enables the process of identifying multiple possible contextual dimensions from the digital script content (during the process of generating an image sequence from the digital script content using GANs), and ensures that the generated image sequence is scaled up or down with respect to the selected dimensions. Similarly, System 100 enables the process of modifying the image sequence (using GANs) based on interactions with the identified dimensions.

[0010] System 100 further enables a process of extracting contextual dimensions from the digital text story content presented with the generated image sequence, thereby enabling scriptwriter entities to modify the dimensions based on determined needs. The modified dimensions enable modification of the image sequence. System 100 can be configured to enable scriptwriter entities to add additional dimensions with respect to the generated image sequence. The newly added dimensions are modified with respect to the generated image sequence, thereby modifying the image sequence and (via the execution of an inverse GAN model) the associated text script content. System 100 may be further configured to enable scriptwriter entities to selectively modify, delete, or add one or more objects to the generated image sequence, or a combination thereof, and to dynamically modify the written text story content.

[0011] The scriptwriter entity can be enabled to split or join multiple image sequences (while interacting with the image sequence), thereby enabling an automated process of splitting or merging text story contexts to create new story content. The virtual reality (VR) user interface can allow users to interact with various image sequences, change the dimensions of the context, and dynamically modify the text story content.

[0012] The system 100 in Figure 1 consists of GAN hardware 139, a text / digital story input component 140, a hardware interface 115, and a network interface controller, all interconnected via a network 117. 153The GAN hardware 139 includes a sensor 112, a circuit 127, and software / hardware 121. The hardware interface may include any type of hardware-based interface, including, in particular, a virtual reality interface. The GAN hardware 139, the text / digital story input component 140, and the hardware interface 115 may each include an embedded device. In this specification, an embedded device is defined as a dedicated device or computer that includes a combination of computer hardware and software (fixed function or programmable) specifically designed to perform a particular function. A programmable embedded computer or device may include a special programming interface. In one embodiment, the GAN hardware 139, the text / digital story input component 140, and the hardware interface 115 may each include a special hardware device that includes special (non-common) hardware and circuitry (i.e., special discrete non-common analog, digital, and logic-based circuits) for performing the processes described with respect to Figures 1 to 6 (independently or in combination). Special discrete non-common analog, digital, and logic-based circuits (e.g., sensor 112, circuit / logic 127, software / hardware 121, etc.) may include uniquely designed components (e.g., special integrated circuits such as application-specific integrated circuits (ASICs) designed solely to perform automated processes for generating and modifying image sequences related to the text content of a digital story and for improving software techniques related to dynamically changing the associated digital text content). Sensor 112 may include any type of internal or external sensor, including, in particular, GPS sensors, Bluetooth beaconing sensors, cell phone detection sensors, Wi-Fi positioning detection sensors, triangulation detection sensors, activity tracking sensors, temperature sensors, ultrasonic sensors, optical sensors, video retrieval devices, humidity sensors, voltage sensors, network traffic sensors, etc.Network 117 may include any type of network, including, in particular, a local area network (LAN), a wide area network (WAN), the internet, and a wireless network.

[0013] System 100 is capable of performing a process (through the execution of a natural language processing model) to perform text analysis of the digital content of the story script. Based on the analysis, System 100 is configured to identify various contextual dimensions within the story content. Similarly, System 100 analyzes a knowledge corpus with respect to various dimensions, in particular, such as weather-related dimensions, event-related dimensions, location-related dimensions, time-related dimensions, dimensions based on physical X, Y, Z positions, and velocity-related dimensions. System 100 is further configured to correct for varying degrees of contextual dimensions, in particular, such as low-degree bad weather versus high-degree bad weather. In addition, System 100 is configured to generate an image sequence from the digital story script content. The image sequence is useful for identifying various possible dimensions from the text story script content.

[0014] Figure 2 shows an algorithm illustrating in detail the process flow enabled by the system 100 in Figure 1 to improve software techniques related to generating and modifying image sequences associated with the text content of a digital story and dynamically changing the associated digital text content, according to embodiments of the present invention. Each step in the algorithm in Figure 2 can be enabled and executed in any order by a computer processor executing computer code. Furthermore, each step in the algorithm in Figure 2 is enabled by the GAN hardware 139, text / digital story input component 140, and may be enabled and executed in combination by the hardware interface 115. In step 200, an image sequence related to the text content of the digital story is generated by the GAN hardware device. In step 202, multiple context dimensions within the text content are identified (through the execution of NLP code). Context dimensions may include, in particular, dimensions such as weather dimension, event dimension, location dimension, time dimension, physical X, Y, Z location dimension, and velocity dimension.

[0015] In step 204, a group of dimensions of multiple contextual dimensions is selected in response to user input. In step 208, the image sequence is scaled up or down in combination with the group of dimensions. In step 210, the image sequence is modified based on the detected interaction with the group of dimensions. In step 212, dimensions are extracted from the group of dimensions during the presentation of the digital story and image sequence. In step 214, a scriptwriter entity (related to the text content of the digital story) is activated to modify the dimensions. In step 216, the image sequence is modified based on the dimensional modifications that occur in response to the results of step 214. In step 218, a hardware interface device is activated to interact with various image sequences of the image sequence and modify multiple contextual dimensions. The hardware interface device may include a virtual reality (VR) interface device. In step 220, the text content of the digital story is dynamically changed.

[0016] In step 224, the script writer entity functionality can be enabled (via GAN hardware) as described in the following implementation scenario.

[0017] In the first scenario, the scriptwriter entity is activated to add an additional contextual dimension to the image sequence (via the hardware interface device) such that the image sequence is modified. Subsequently, the text content is modified by running the inverse GAN model with respect to the result of the modification of the image sequence.

[0018] The second scenario activates the scriptwriter entity to selectively modify at least one visual object of the image sequence such that the text content is modified via running the inverse GAN model with respect to the result of activating the scriptwriter entity (via the hardware interface device).

[0019] The third scenario activates the scriptwriter entity to selectively remove at least one visual object from the image sequence such that the text content is modified via running the inverse GAN model with respect to the result of activating the scriptwriter entity (via the hardware interface device).

[0020] The fourth scenario activates the scriptwriter entity to selectively add at least one visual object to the image sequence such that the text content is modified via running the inverse GAN model with respect to the result of activating the scriptwriter entity (via the hardware interface device).

[0021] The fifth scenario activates the scriptwriter entity to split a plurality of image sequences of the image sequence (via the hardware interface device during interaction with the various image sequences). In response, the text content is split and new text content is generated for the digital story.

[0022] The sixth scenario involves enabling a script writer entity to combine multiple image sequences (via a hardware interface device during interaction with the various image sequences). In response, text content is merged, and new text content for the digital story is generated.

[0023] Figure 3 is an internal structure diagram of the software / hardware 121 (i.e., 121) of Figure 1 according to an embodiment of the present invention. The software / hardware 121 includes an identification module 304, a modification module 305, an extraction module 308, a modification / activation module 314, and a communication controller 312. The identification module 304 includes dedicated hardware and software for controlling all functions related to the identification step in Figure 2. The modification module 305 includes dedicated hardware and software for controlling all functions related to the modification step described with respect to the algorithm in Figure 2. The extraction module 308 includes dedicated hardware and software for controlling all functions related to the extraction step in Figure 2. The modification / activation module 314 includes dedicated hardware and software for controlling all functions related to the modification and activation steps of the algorithm in Figure 2. The communication controller 312 is activated to control all communication between the identification module 304, the modification module 305, the extraction module 308, and the modification / activation module 314.

[0024] Figure 4 shows a system 400 according to an embodiment of the present invention, which includes a GAN module 402a and an NLP module 404 for modifying a digital script 405 of digital story content. The system 400 is configured to present multiple possible dimensions 412 associated with an image sequence 408 generated from the digital script 405, thereby enabling a scriptwriter entity to modify the dimensions 412. Similarly, the GAN module 402a is configured to modify the images in the generated image sequence 408 so that the digital story content is dynamically updated. Modifying the images in the generated image sequence 408 results in the generation of a modified image sequence 410. The GAN module 402a may be enabled to generate an image sequence corresponding to input text, and the resulting generated images may be viewed by a user.

[0025] NLP module 404 may be configured to search a knowledge corpus containing input text and various dimensions 412 of a digital script 405 (e.g., weather dimension, color dimension, etc.). In response, system 400 analyzes the input text to identify the dimensions available within the digital script 405. Various dimensions 412 of the input text and their relative degrees are identified. For example, the "weather" dimension may include various degrees associated with it, such as sunny, cloudy, windy, rainy, etc. The identified dimensions 412 and their relative degrees may be displayed for the user. Similarly, the user may modify (e.g., add, update, delete, etc.) the dimensions 412 and their degrees with respect to selection. The image generated (from GAN module 402a), and the modified dimensions and associated dimensional degrees selected by the user, are sent to a second GAN module 402b as conditional input. System 400 is further configured to execute a conditional text-to-image translation code through the use of a cycle-consistent adversarial network that searches for input (i.e., user-selected images and modified dimensions). The text-to-image module (GAN module 402b) is enabled to generate a modified version of the input image with respect to user-selected dimensions and degree. The inverse GAN module 402c (i.e., the image-to-text translation module) takes a modified image sequence 410 as input and generates the associated text (i.e., modified script 415). The modified image sequence 410 and the corresponding modified script 415 are used by the user to finalize the digital script.

[0026] Figures 5A to 5D illustrate a process 500 for modifying a digital script 502 and generating corresponding image sequences 504a and 504b according to an embodiment of the present invention. Process 500 is initiated when the text content (i.e., the digital script 502) is provided as input to the text-to-image GAN module 506 and NLP module 507 for performing text analysis of the digital script 502 (with respect to the knowledge corpus 509). In response to the text analysis of the digital script 502, the system 500 identifies various context (and degree dimensions) 511 from the story content of the digital script 502. Context dimensions 511 may include, in particular, weather dimensions, event dimensions, location dimensions, time dimensions, physical X, Y, Z location dimensions, velocity dimensions, etc. The system 500 further enables various degrees of the context dimensions 511. Subsequently, the GAN module 506 generates an image sequence 504a from the text content of the digital script 502 (through the execution of the GAN module 506) and identifies various possible dimensions (of the context dimension 511) from the text content. The image sequence 504a may be displayed via a hardware / software interface 514 (e.g., a 2D display, a VR device, etc.). The system 500 then presents one or more context dimensions in combination with the image sequence 504a so that the story script writer entity 517 can be enabled (through the text-to-image GAN network component 522 and the image-to-text GAN network component 524) to modify the degree of the various dimensions presented with the image sequence 504a. In response to the selection of various modified context dimensions 519, the system 500 receives the relevant user input and identifies the modified context dimension 519. The modified context dimension 519 is used to analyze the current image sequence (i.e., image sequence 504a) in order to modify the images within the image sequence 504a. All modifications to context dimension 511 are taken into consideration to update the images in image sequence 504a.Similarly, system 500 activates scriptwriter entity 517 to add additional dimensions and selected dimensional degrees to image sequence 504a, and the images in image sequence 504a are updated accordingly. Scriptwriter entity 517 may be activated to selectively modify / add / remove one or more image objects from image sequence 504a so that the images in image sequence 504a are changed. Scriptwriter entity 517 can split or join different images, resulting in an updated image sequence 504b, which can be viewed via hardware-software interface 527. Once the modification process is complete, system 400 performs the process of updating the digital script 502 with the modified images of image sequence 504b, resulting in the generation of a modified digital script 528.

[0027] Figure 6 is a detailed diagram of the text-to-image GAN network component 522 of Figure 5, according to an embodiment of the present invention. The GAN network component 522 includes a first stage 602 (stage 1) and a second stage 604 (stage 2). The first stage 602 includes a pair of generator G1 and discriminator D1. Similarly, the second stage 604 includes a pair of generator G2 and discriminator D2. Generator G1 is configured to produce a low-resolution image 607 (e.g., 64x64ppi), and generator G2 is configured to produce a high-resolution image 609 (128x128ppi). Relevant text embedding data 605 (i.e., script) and associated noise can be used as input to the first stage 602. Furthermore, the image and user-modified dimensions and degree may also be used as input to the first stage 602. Generator G1 may be configured to skip the thought text embedding and produce a composite image (i.e., the low-resolution image 607). Similarly, the first stage 602 classifier D1 is conditioned on the same text embeddings and trained to classify between real and composite images with a resolution of 64x64ppi. The second stage generator G1 includes a series of upsampling blocks 611. The upsampling blocks 611 include enabling a nearest-neighbor upsampling process followed by a 3x3 stride 1 convolution process to project the input onto a 3x64x64 image (i.e., the low-resolution image 607) containing the low-resolution (64x64ppi) image. The classifier D1 includes a series of downsampling blocks 612 that project the input onto a 512x4x4 dimension. The aforementioned 512x4x4 dimension is concatenated with a 128-dimensional compressed embedding and uses a sigmoid layer 615 to produce an output between 0 (fake) and 1 (real) to distinguish the low-resolution image. The second stage generator G1 takes I1 along with the embeddings as input and generates a higher-resolution 128x128 image.

[0028] The generator G2 includes a series of downsampling blocks 622 that project the 3x64x64 input image 624 to a 512x16x16 dimension. A 128-dimensional embedding is then concatenated. The input image 624 is transmitted as a series of residual blocks 626 followed by a series of upsampling blocks 628 to generate an image 609 (i.e., a 128x128 image) (higher resolution). The discriminator D2 receives the image 609 (as input). The discriminator D2 includes a series of downsampling blocks 629, allowing the sigmoid layer 632 to generate an output between 0 (fake) and 1 (real) to distinguish the high resolution image (128x128). Thus, the output of the first stage 602 is used as input to the second stage 604 to generate a higher resolution image (i.e., high resolution image 609) as an output containing the user-modified dimension.

[0029] Figure 7 shows a detailed diagram of the image-to-text GAN network component 524 (i.e., the caption GAN network component) of Figure 5, according to an embodiment of the present invention. The GAN network component 524 includes components 702 and 704. Component 702 forms a caption generator that includes convolutional neural network (CNN) features 708 and long-term short-term memory (LSTM) components 707a...707n for retrieving noise Z709 as input for outputting a caption 711. The input to component 702 includes the output high-resolution image 715 (from the GAN network component 522 in Figure 6), including user-modified dimensions and degree.

[0030] Component 704 forms a discriminator by performing a dot product with respect to the CNN features 712 of the high-resolution modified image 715a and the outputs from the LSTM components 707a...707n. The high-resolution modified image 715a, containing the user's preferred dimensions and degrees, is sent as input to perform a regular sequence modeling process with respect to (LSTM components 717a...717n) to generate the corresponding. The resulting output text script 720 contains the user-updated dimensions and degrees, which are available for the user to make a final decision.

[0031] Figure 8 shows a computer system 90 (e.g., GAN hardware 139 in Figure 1, text / digital story input component) used by or configured by the system 100 in Figure 1 to improve software techniques related to generating and modifying image sequences associated with the text content of a digital story and dynamically changing the associated digital text content, according to embodiments of the present invention. 140 This shows the hardware interface 115).

[0032] Aspects of the present invention may take the form of entirely hardware embodiments, entirely software embodiments (including firmware, resident software, microcode, etc.), or embodiments combining software and hardware embodiments, all of which are generally referred to herein as “circuits,” “modules,” or “systems.”

[0033] The present invention may be a system, a method, a computer program product, or a combination thereof. The computer program product may include a computer-readable storage medium storing computer-readable program instructions for causing a processor to execute an aspect of the present invention.

[0034] A computer-readable storage medium can be a tangible device capable of holding and storing instructions used by an instruction execution device. A computer-readable storage medium may, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or a suitable combination thereof. More specific examples of computer-readable storage media include portable computer diskettes, hard disks, RAM, ROM, EPROM or flash memory, SRAM, CD-ROM, DVD, memory stick, floppy disk, punch cards or grooved raised structures, and mechanically encoded devices on which instructions are recorded, and suitable combinations thereof. As used herein, computer-readable storage media should not be interpreted as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses passing through optical fiber cables), or electrical signals transmitted through wires.

[0035] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof). The network consists of copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. The network adapter card or network interface of each computing / processing device receives computer-readable program instructions from the network and transfers the computer-readable program instructions for storage on the computer-readable storage medium within each computing / processing device.

[0036] The computer-readable program instructions for performing the operation of the present invention may be assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, Spark, and R, and procedural programming languages ​​such as the C programming language or similar programming languages. The computer-readable program instructions are executable as a standalone software package, either entirely on the user's computer or partially on the user's computer. Alternatively, they may be executable partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or wide area network (WAN), or to an external computer (for example, via the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA) can execute computer-readable program instructions by personalizing them using state information of computer-readable program instructions in order to perform aspects of the present invention.

[0037] Aspects of the present invention are described herein with reference to flowcharts or block diagrams, or both, of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It will be understood that each block in a flowchart or block diagram, or both, and any combination of blocks in a flowchart or block diagram, or both, can be implemented by computer-readable program instructions.

[0038] These computer-readable program instructions can be provided to a general-purpose computer, a dedicated computer processor, or other programmable data processing device to generate a machine, such that instructions executed via the processor of a computer or other programmable data processing device generate means for implementing functions / operations specified in one or more blocks of a flowchart or block diagram or both. These computer-readable program instructions can also be stored in a computer-readable storage medium that can be connected to a computer, a programmable data processing device, or other device or combination of devices that function in a particular way, such that the computer-readable storage medium on which the instructions are stored constitutes one of the outputs containing instructions that implement the modes of functions / operations specified in one or more blocks of a flowchart or block diagram or both.

[0039] Computer-readable program instructions, like instructions that perform a function / action specified in one or more blocks of a flowchart or block diagram or both on a computer, other programmable device, or other device, can also be loaded into a computer, other programmable data processing device, or other device and perform a series of operational steps on the computer, other programmable device, or other device to produce a computer-implemented process.

[0040] The flowcharts and block diagrams in the figures illustrate the configuration, function, and operation of executable implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or part of an instruction, which constitutes one or more executable instructions for implementing a specified logical function. In some alternative embodiments, the functions shown in the blocks may differ from the order shown in the figures. For example, two blocks shown consecutively may actually be achieved as a single step, executed simultaneously, substantially simultaneously, partially or entirely in overlapping time, or the blocks may be executed in reverse order depending on the functions involved. It should also be noted that each block in a block diagram or flowchart diagram, or both, and any combination of blocks in a block diagram or flowchart diagram, or both, can be implemented by a special-purpose hardware-based system that performs a specified function or operation, or a combination of special-purpose hardware and computer instructions.

[0041] The computer system 90 shown in Figure 8 includes a processor 91, an input device 92 coupled to the processor 91, an output device 93 coupled to the processor 91, and storage devices 94 and 95 coupled to the processor 91, respectively. The input device 92 may be, in particular, a keyboard, mouse, camera, touchscreen, etc. The output device 93 may be, in particular, a printer, plotter, computer screen, magnetic tape, removable hard disk, floppy disk, etc. The storage devices 94 and 95 may be, in particular, optical storage devices such as hard disks, floppy disks, magnetic tape, compact discs (CDs) or digital video discs (DVDs), dynamic random access memory (DRAM), read-only memory (ROM), etc. The storage device 95 includes computer code 97. The computer code 97 includes algorithms (e.g., the algorithm in Figure 2) for improving software techniques related to generating and modifying image sequences related to the text content of a digital story and dynamically changing the related digital text content. The processor 91 executes the computer code 97. The storage device 94 includes input data 96. Input data 96 includes inputs requested by computer code 97. Output device 93 displays output from computer code 97. Either or both of storage devices 94 and 95 (or one or more additional storage devices such as read-only storage device 85) can be used as computer-readable media (or computer-readable media or program storage device) having an algorithm (e.g., the algorithm in Figure 2) and having computer-readable program code implemented therein, or other data stored therein, or both, the computer-readable program code includes computer code 97. Generally, a computer program product (or, alternatively, a manufactured product) of a computer system 90 may include computer-readable media (or program storage devices).

[0042] In some embodiments, rather than being stored and accessed from a hard drive, optical disc, or other writable, rewritable, or removable hardware storage device 95, stored computer program code 84 (e.g., including algorithms) may be stored in a static, non-removable, read-only storage medium such as a read-only storage (ROM) device 85, or may be accessed directly by the processor 91 from such a static, non-removable, read-only medium. Similarly, in some embodiments, stored computer program code 97 may be stored as computer-readable firmware 85, or accessed directly by the processor 91 from such firmware 85, rather than from a more dynamic or removable hardware data storage device 95 such as a hard drive or optical disc.

[0043] Nevertheless, any component of the present invention may be created, integrated, hosted, maintained, deployed, managed, serviced, etc. by a service supplier that provides services to improve software techniques related to generating and modifying image sequences associated with the text content of a digital story and dynamically changing the associated digital text content. Accordingly, the present invention discloses a process for deploying, creating, integrating, hosting, maintaining, or integrating, or a combination thereof, a computing infrastructure, including integrating computer-readable code into a computer system 90, the code combined with the computer system 90, can perform methods to enable the process of improving software techniques related to generating and modifying image sequences associated with the text content of a digital story and dynamically changing the associated digital text content. In another embodiment, the present invention provides a business method for performing the process steps of the present invention on a subscription, advertising, or fee basis, or a combination thereof. That is, a service supplier, such as a solution integrator, may provide services to enable the process of improving software techniques related to generating and modifying image sequences associated with the text content of a digital story and dynamically changing the associated digital text content. In this case, the service supplier may create, maintain, support, etc., a computer infrastructure that performs the process steps of the present invention for one or more customers. In return, the service supplier may receive payments from customers based on a subscription or fee agreement, or both, or from the sale of advertising content to one or more third parties, or both.

[0044] Figure 8 shows computer system 90 as a specific hardware and software configuration, but any hardware and software configuration known to those skilled in the art can be used in combination with the specific computer system 90 in Figure 8 for the purposes described above. For example, storage devices 94 and 95 may be part of a single storage device rather than being separate storage devices.

[0045] <Cloud Computing Environment> This disclosure includes a detailed description of cloud computing, but the implementations of the teachings described herein are not limited to cloud computing environments. Rather, embodiments of the present invention can be implemented in any other type of computer environment that is currently known or may be developed in the future.

[0046] Cloud computing is a service delivery model that enables convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal administrative effort or interaction with service providers. This cloud model may include at least five characteristics, at least three service models, and at least four implementation models.

[0047] The characteristics are as follows:

[0048] On-demand self-service: Cloud consumers can unilaterally prepare computing power, such as server time and network storage, automatically as needed, without requiring human interaction with service providers.

[0049] Broad network access: Computing power is available over the network and accessible through standard mechanisms. This facilitates utilization by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, PDAs).

[0050] Resource pooling: A provider's computing resources are pooled and delivered to multiple consumers using a multi-tenant model. Various physical and virtual resources are dynamically allocated and reallocated as needed. Generally, consumers have a sense of location independence because they do not manage or know the exact location of the resources provided. However, consumers may be able to identify the location at a higher level of abstraction (e.g., country, state, data center).

[0051] Rapid Elasticity: Computing power can be prepared quickly and flexibly, allowing it to scale out automatically and immediately, and to be quickly released and scale in immediately. To consumers, the computing power available for preparation often appears unlimited and can be purchased in any quantity at any time.

[0052] Measured Services: Cloud systems leverage metric capabilities at a certain level of abstraction, appropriate for the type of service (e.g., storage, processing, bandwidth, active user accounts), to automatically control and optimize resource usage. Resource usage can be monitored, controlled, and reported, providing transparency to both service providers and consumers.

[0053] The service model is as follows:

[0054] Software as a Service (SaaS): The functionality offered to consumers is the ability to use the provider's applications running on a cloud infrastructure. These applications can be accessed from various client devices via thin client interfaces such as web browsers (e.g., webmail). Consumers do not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or even individual application functions, except for configuring a limited number of user-specific applications.

[0055] Platform as a Service (PaaS): The functionality offered to consumers is the ability to deploy applications they have created or acquired to cloud infrastructure using programming languages ​​and tools supported by the provider. Consumers do not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, and storage, but they can control the deployed applications and, in some cases, the configuration of their hosting environment.

[0056] Infrastructure as a Service (IaaS): The functionality provided to consumers is the provision of processors, storage, networking, and other basic computing resources that enable consumers to deploy and run any software, including operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but they can control the operating system, storage, and deployed applications, and in some cases, partially control certain network components (e.g., host firewalls).

[0057] The deployment model is as follows:

[0058] Private Cloud: This cloud infrastructure is operated exclusively for a specific organization. This cloud infrastructure can be managed by that organization or a third party and can reside on-premises or off-premises.

[0059] Community Cloud: This cloud infrastructure is shared by multiple organizations to support a specific community with common interests (e.g., mission, security requirements, policies, and compliance). This cloud infrastructure can be managed by the organization or a third party and can reside on-premises or off-premises.

[0060] Public Cloud: This cloud infrastructure is provided to a large number of people or large industry groups and is owned by organizations that sell cloud services.

[0061] Hybrid Cloud: This cloud infrastructure combines two or more cloud models (private, community, or public). While maintaining the unique entities of each model, they are bound together by standards or individual technologies to achieve data and application portability (e.g., cloud bursting for load balancing across clouds).

[0062] Cloud computing environments are service-oriented environments that emphasize statelessness, low coupling, modularity, and semantic interoperability. At the core of cloud computing is the infrastructure, which includes a network of interconnected nodes.

[0063] Referring to Figure 9, an exemplary cloud computing environment 50 is shown. As illustrated, the cloud computing environment 50 includes one or more cloud computing nodes 10. Local computer devices used by cloud consumers (e.g., personal digital assistants (PDAs) or mobile phones 54A, desktop computers 54B, laptop computers 54C, or automotive computer systems 54N, or a combination thereof) can communicate with these nodes. The nodes 10 can communicate with each other. The nodes 10 can be grouped physically or virtually (not shown) in one or more networks, such as the private, community, public, or hybrid clouds or a combination thereof. This allows the cloud computing environment 50 to provide infrastructure, platforms, or software as a service, or a combination thereof, without requiring cloud consumers to maintain resources on their local computer devices. Note that the types of computer devices 54A, 54B, 54C, and 54N shown in Figure 9 are illustrative only, and it should be understood that the computing nodes 10 and the cloud computing environment 50 can communicate with any type of electronic device via any type of network or network addressable connection (e.g., using a web browser) or both.

[0064] Referring to Figure 10, a set of functional abstraction layers provided by the cloud computing environment 50 (Figure 9) is shown. It should be understood that the components, layers, and functions shown in Figure 10 are illustrative only, and embodiments of the present invention are not limited to these. As illustrated, the following layers and corresponding functions are provided.

[0065] The hardware and software layer 60 includes hardware components and software components. Examples of hardware components include a mainframe 61, a reduced instruction set computer (RISC) architecture-based server 62, server 63, blade server 64, storage 65, and a network and network components 66. In some embodiments, the software components include network application server software 67 and database software 68.

[0066] The virtualization layer 70 provides an abstraction layer. From this layer, for example, the following virtual entities can be provided: virtual servers 71, virtual storage 72, virtual networks 73 including virtual private networks, virtual applications and operating systems 74, and virtual clients 75.

[0067] As an example, the management layer 80 can provide the following functions: Resource preparation 81 enables the dynamic procurement of computing resources and other resources used to perform tasks within the cloud computing environment. Metering and pricing 82 enables cost tracking as resources are used within the cloud computing environment and billing or invoicing for the consumption of these resources. As an example, these resources may include licenses for application software. Security enables not only protection of data and other resources, but also identification and verification of cloud consumers and tasks. The user portal 83 provides consumers and system administrators with access to the cloud computing environment. Service level management 87 enables the allocation and management of cloud computing resources to ensure that requested service levels are met. Service Level Agreement (SLA) planning and execution 88 enables the pre-arrangement and procurement of cloud computing resources that are expected to be needed in the future in accordance with the SLA.

[0068] Workload layer 101 provides examples of the capabilities available to the cloud computing environment. Examples of workloads and capabilities available from this layer include mapping and navigation 102, software development and lifecycle management 103, delivery of virtual classroom education 133, data analysis processing 134, transaction processing 106, and improvements to software technologies related to generating and modifying image sequences associated with the text content of digital stories, and dynamic modification of associated digital text content 107.

[0069] While embodiments of the present invention have been described herein for illustrative purposes, many modifications and changes will become apparent to those skilled in the art. Therefore, the appended claims are intended to encompass all such modifications and changes that fall within the scope of the present invention.

Claims

1. A generative adversarial network (GAN) hardware device including a processor coupled to a computer-readable memory unit, wherein the memory unit includes instructions that, when executed by the processor, perform a digital script modification method that enables natural language processing (NLP), and the method is The aforementioned processor generates an image sequence related to the text content of the digital story, The processor identifies multiple contextual dimensions within the text content by executing NLP code, The processor selects a group of dimensions of the plurality of context dimensions in response to user input, The processor enlarges or reduces the image sequence by combining it with the group of dimensions, The processor modifies the image sequence based on the detected interaction with the group of dimensions, The processor extracts dimensions from the group of dimensions during the presentation of the digital story and the image sequence, The processor enables the scriptwriter associated with the text content of the digital story in order to correct the dimension, The processor modifies the image sequence based on the dimensional correction that occurs in response to the activation. The processor enables a hardware interface device to interact with various image sequences of the image sequence and to change the multiple context dimensions. The processor dynamically changes the text content of the digital story in response to the activation, Generative Adversarial Network (GAN) hardware devices, including those mentioned above.

2. The GAN hardware device according to claim 1, wherein the plurality of context dimensions include dimensions selected from the group consisting of weather dimensions, event dimensions, location dimensions, time dimensions, physical X, Y, Z location dimensions, and velocity dimensions.

3. The above method further, The processor enables the script writer via the hardware interface device to add an additional contextual dimension to the image sequence, The processor performs a first modification to the image sequence with respect to the additional context dimension, The processor, which executes an inverse GAN model with respect to the result of the first modification, performs a second modification on the text content. A GANG hardware device according to claim 1, including the above.

4. The above method further, The processor enables the script writer via the hardware interface device to selectively modify at least one visual object in the image sequence, The text content is modified by the processor that executes an inverse GAN model with respect to the result of the script writer being activated, A GANG hardware device according to claim 1, including the above.

5. The above method further, The processor enables the script writer via the hardware interface device to selectively remove at least one visual object from the image sequence, The text content is modified by the processor that executes an inverse GAN model with respect to the result of the script writer being activated, A GANG hardware device according to claim 1, including the above.

6. The above method further, The processor enables the script writer via the hardware interface device to selectively add at least one visual object to the image sequence, The text content is modified by the processor that executes an inverse GAN model with respect to the result of the script writer being activated, A GANG hardware device according to claim 1, including the above.

7. The above method further, The processor enables the script writer to divide the image sequence into multiple image sequences via the hardware interface device during interaction with the various image sequences, The processor divides the text content in response to the result of enabling the script writer, The processor generates new text content for the digital story in response to the division, A GANG hardware device according to claim 1, including the above.

8. The above method further, The processor enables the script writer via the hardware interface device to concatenate multiple image sequences of the image sequence during interaction with the various image sequences, The processor, in response to the activation result, merges the text content, The processor generates new text content for the digital story in response to the merge, A GANG hardware device according to claim 1, including the above.

9. The GAN hardware device according to claim 1, wherein the hardware interface device comprises a virtual reality (VR) interface device.

10. A digital script modification method that enables natural language processing (NLP), The processor of a Generative Adversarial Network (GAN) hardware device generates image sequences related to the text content of a digital story, The processor identifies multiple contextual dimensions within the text content by executing NLP code, The processor selects a group of dimensions of the plurality of context dimensions in response to user input, The processor enlarges or reduces the image sequence by combining it with the group of dimensions, The processor modifies the image sequence based on the detected interaction with the group of dimensions, The processor extracts dimensions from the group of dimensions during the presentation of the digital story and the image sequence, The processor enables the scriptwriter associated with the text content of the digital story in order to correct the dimension, The processor modifies the image sequence based on the dimensional correction that occurs in response to the activation. The processor enables a hardware interface device to interact with various image sequences of the image sequence and to change the multiple context dimensions. The processor dynamically changes the text content of the digital story in response to the activation, Methods that include...

11. The method according to claim 10, wherein the plurality of context dimensions include dimensions selected from the group consisting of weather dimensions, event dimensions, location dimensions, time dimensions, physical X, Y, Z location dimensions, and velocity dimensions.

12. The above method further, The processor enables the script writer via the hardware interface device to add an additional contextual dimension to the image sequence, The processor performs a first modification to the image sequence with respect to the additional context dimension, The processor, which executes an inverse GAN model with respect to the result of the first modification, performs a second modification on the text content. The method according to claim 10, including the method described in claim 10.

13. The processor enables the script writer via the hardware interface device to selectively modify at least one visual object in the image sequence, The text content is modified by the processor that executes an inverse GAN model with respect to the result of the script writer being activated, The method according to claim 10, further comprising:

14. The processor enables the script writer via the hardware interface device to selectively remove at least one visual object from the image sequence, The text content is modified by the processor that executes an inverse GAN model with respect to the result of the script writer being activated, The method according to claim 10, further comprising:

15. The processor enables the script writer via the hardware interface device to selectively add at least one visual object to the image sequence, The text content is modified by the processor that executes an inverse GAN model with respect to the result of the script writer being activated, The method according to claim 10, further comprising:

16. The processor enables the script writer to divide the image sequence into multiple image sequences via the hardware interface device during interaction with the various image sequences, The processor divides the text content in response to the result of enabling the script writer, The processor generates new text content for the digital story in response to the division, The method according to claim 10, further comprising:

17. The processor enables the script writer via the hardware interface device to concatenate multiple image sequences of the image sequence during interaction with the various image sequences, The processor, in response to the activation result, merges the text content, The processor generates new text content for the digital story in response to the merge, The method according to claim 10, further comprising:

18. The method according to claim 10, wherein the hardware interface device comprises a virtual reality (VR) interface device.

19. A computer system provides at least one support service for at least one of the creation, integration, hosting, maintenance, and deployment of computer-readable code, wherein the computer-readable code is executed by the processor and provides the processor to perform the generation, identification, selection, scaling or reduction, modification, extraction, activation of the script writer, modification, activation of the hardware interface device, and dynamic modification. The method according to claim 10, further comprising:

20. The computer-readable program code includes an algorithm that performs a digital script modification method that enables natural language processing (NLP) when executed by the server's processor, and the method is The aforementioned processor generates an image sequence related to the text content of the digital story, The processor identifies multiple contextual dimensions within the text content by executing NLP code, The processor selects a group of dimensions of the plurality of context dimensions in response to user input, The processor enlarges or reduces the image sequence by combining it with the group of dimensions, The processor modifies the image sequence based on the detected interaction with the group of dimensions, The processor extracts dimensions from the group of dimensions during the presentation of the digital story and the image sequence, The processor enables the scriptwriter associated with the text content of the digital story in order to correct the dimension, The processor modifies the image sequence based on the dimensional correction that occurs in response to the activation. The processor enables a hardware interface device to interact with various image sequences of the image sequence and to change the multiple context dimensions. The processor dynamically changes the text content of the digital story in response to the activation, A computer program that includes [this].

Citation Information

Patent Citations

  • Video service providing method and service server using the same

    JP2019212308A

  • Artificial intelligence in interactive storytelling

    US20190304157A1