System and method for conversion of text-based content to consumption formats
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2026-02-10
- Publication Date
- 2026-08-13
Smart Images

Figure US2026014765_13082026_PF_FP_ABST
Abstract
Description
Patent Docket No. 140276-5001SYSTEM AND METHOD FOR CONVERSION OF TEXT-BASED CONTENT TO CONSUMPTION FORMATS Cross-Reference to Related Applications
[0001] This application claims priority to U.S. Provisional Application No. 63 / 756,739, filed on February 10, 2025 and entitled “System And Method For Conversion of Text-Based Content To Consumption Formats,” the contents of which are hereby incorporated herein in their entirety.Technical Field
[0002] This application relates generally to conversion of content to different consumable formats, and more particularly, to the use of artificial intelligence models to automate the conversion of text based content to multiple consumption formats to be chosen, at consumption time, by a consumer.Background
[0003] People and entities that publish text-based content, e.g. web pages, struggle to craft content that reaches the right audience, at the right time, in the right format. Different segments of the consuming audience may wish to consume content in different formats. Some may prefer long form text, while others may prefer short form text, graphic formats containing information, audio, or video. However, creating the same content, manually, in each of these different formats, is costly.
[0004] It would be advantageous for there to be an automated system for transforming a single unit of text content into different consumption formats to reach different audience segments.Summary
[0005] A method is herein disclosed. The method includes the steps of transmitting, to a format conversion server, a request to convert a unit of text based content to at least two consumption formats. At the format conversion server, the method includes transmitting the unit of text based content, and a first artificial intelligence (“Al”) configuration parameter, to a first Al model, the first Al model being configured to convert the text based content to a first converted content in a first consumption format, transmitting the unit text based content, and a second Al configuration parameter, to a second artificial intelligence model, the second Al model beingPatent Docket No. 140276-5001configured to convert the text based content to a second converted content in a second consumption format, wherein the first consumption format and the second consumption format are each one of graphic, video, audio, and summarized text, and wherein the first consumption format is separate and distinct from the second consumption format, and storing the substantially text based content, the first converted content, and the second converted content.
[0006] The method further includes transmitting data representing a user interface to a user device, the user interface including the text based content and a format selector interface element. The method further includes displaying the text based content and the format selector interface element on the user device. The method further includes, in response to a user interaction with a first portion of the format selector interface element, said first portion indicating the first consumption format, transmitting data representing the first converted content to the user device, and displaying the first converted content on the user device; and, in response to a user interaction with a second portion of the format selector interface element, said second portion indicating the second consumption format, transmitting data representing the second converted content to the user device, and displaying the second converted content on the user device.
[0007] In some embodiments, the first Al configuration parameter and the second Al configuration parameter include editor preferences, received via user input prior to transmitting the request to the conversion server, the editor preferences for the first prompt are specific to the unit of text based content being transmitted and contain presentation parameters for the first consumption format; and the editor preferences for the second prompt are specific to the unit of text based content being transmitted and contain presentation parameters for the second consumption format.
[0008] In some embodiments, the first consumption format is a video, and the presentation parameters are from the group consisting of video orientation, video duration, voice-over voice, voice-over pitch, and voice-over speed. In some embodiments, converting the unit of text based content into a first converted content includes creating a summary of the text based content, wherein a length of the summary is configured to match the video duration.
[0009] In some embodiments, the first consumption format is audio, and the presentation parameters are from the group consisting of duration, voice, voice pitch, and speed of speech. InPatent Docket No. 140276-5001some embodiments, converting the substantially text based content into a first converted content includes creating a summary of the unit of text based content, wherein the summary is configured to match the duration.
[0010] In some embodiments, the first consumption format is summarized text, wherein the presentation parameters include a text length, and wherein converting the substantially text based content into a first converted content includes creating a summary of the unit of text based content, wherein the summary is configured to match the length.
[0011] In some embodiments, the first consumption format is a graphic containing text, wherein the text contained in the graphic is a summary of the unit of text based content, generated by the first Al, and designed to fit on the graphic. In some embodiments, the graphic further contains at least one chart generated by the first Al model based on the unit of text based content.
[0012] In some embodiments, the first consumption format is summarized text, and wherein the first Al configuration parameter includes a specified length for the summarized text.Brief Description of the Drawings
[0013] The features and advantages of the present invention will be more fully disclosed in, or rendered obvious by the following detailed description of the preferred embodiments, which are to be considered together with the accompanying drawings wherein like numbers refer to like parts and further wherein:
[0014] FIG. 1 illustrates a system architecture diagram of a system configured to provide conversion of text-based content to consumption formats, in accordance with some embodiments;
[0015] FIG. 2 illustrates a computer system configured to implement one or more processes, in accordance with some embodiments;
[0016] FIG. 3 A through 3C illustrate user interface examples of various aspects of a system for conversion of text-based content to consumption formats;
[0017] FIGs. 4A and 4B is a process flow illustrating steps of a method for conversion of text-based content to consumption formats, in accordance with some embodiments.Patent Docket No. 140276-5001
[0018] FIG. 5A illustrates an artificial neural network, which may be utilized by at least some systems and methods described herein, in accordance with some embodiments;
[0019] FIG. 5B illustrates a tree-based artificial neural network, which may be utilized by at least some systems and methods described herein, in accordance with some embodiments;
[0020] FIG. 5C illustrates a deep neural network (DNN), which may be utilized by at least some systems and methods described herein, in accordance with some embodiments;
[0021] FIG. 6A is a flowchart illustrating a training method for generating a trained machine learning model, in accordance with some embodiments; and
[0022] FIG. 6B is a process flow illustrating various steps of the training method of FIG.6 A, in accordance with some embodiments.
[0023] FIG. 7 is a structural diagram of an LLM formed in a transformer architecture, which may be utilized by at least some systems and methods described herein, in accordance with some embodiments.Detailed Description
[0024] This description of the exemplary embodiments is intended to be read in connection with the accompanying drawings, which are to be considered part of the entire written description. Terms concerning data connections, coupling and the like, such as “connected” and “interconnected,” and / or “in signal communication with” refer to a relationship wherein systems or elements are electrically connected (e.g., wired, wireless, etc.) to one another either directly or indirectly through intervening systems, unless expressly described otherwise. The term “operatively coupled” is such a coupling or connection that allows the pertinent structures to operate as intended by virtue of that relationship.
[0025] In the following, various embodiments are described with respect to the claimed systems as well as with respect to the claimed methods. Features, advantages, or alternative embodiments herein may be assigned to the other claimed objects and vice versa. In other words, claims for the systems may be improved with features described or claimed in the context of the methods. In this case, the functional features of the method are embodied by objective units of thePatent Docket No. 140276-5001systems. While the present disclosure is susceptible to various modifications and alternative forms, specific embodiments are shown by way of example in the drawings and will be described in detail herein. The objectives and advantages of the claimed subject matter will become more apparent from the following detailed description of these exemplary embodiments in connection with the accompanying drawings.
[0026] In some embodiments, systems, and methods for conversion of text-based content to consumption formats includes one or more trained artificial intelligence (“Al”) models. The trained models may include one or more models, such as models configured to summarize text, models configured to convert text to speech, models configured to create a video including text to speech voice over, and models configured to create informational graphics, or infographics, based on and / or including summaries of text.
[0027] In general, a trained function mimics cognitive functions that humans associate with other human minds. In particular, by training based on training data the trained function is able to adapt to new circumstances and to detect and extrapolate patterns.
[0028] In general, parameters of a trained function may be adapted by means of training. In particular, a combination of supervised training, semi-supervised training, unsupervised training, reinforcement learning and / or active learning may be used. Furthermore, representation learning (an alternative term is “feature learning”) may be used. In particular, the parameters of the trained functions may be adapted iteratively by several steps of training.
[0029] In some embodiments, a trained function may include a neural network, a support vector machine, a decision tree, a Bayesian network, a clustering network, Qlearning, genetic algorithms and / or association rules, and / or any other suitable artificial intelligence architecture. In some embodiments, a neural network may be a deep neural network, a convolutional neural network, a convolutional deep neural network, etc. Furthermore, a neural network may be an adversarial network, a deep adversarial network, a generative adversarial network, etc.
[0030] FIG. 1 illustrates a system architecture diagram of a system 10 configured to provide conversion of text-based content to consumption formats, in accordance with some embodiments. In some embodiments, a format selector interface element, such as a toolbar, mayPatent Docket No. 140276-5001be presented along with text content, such as a web page, which would allow the viewer of the text content to, additionally or alternatively, consume the content in a different format, e.g. an informative graphic, a spoken word audio version of the content, a video of the content with an audio voice-over, or a shortened text summary of the full text content. Persons having skill in the art will realize that the phrase “text content,” as used herein, may include content that contains both text and non-textual elements.
[0031] Client applications, running on user devices, such as a mobile device 12, capable of installing mobile applications and taking photographs, a web browser 14 capable of accessing webpages with source content, or an embed toolbar 16, may access a front end of a web application 18. Certain use cases are presented herein, which are exemplary only and not intended to be limiting. In some use cases, an editor may log in via browser 14. “Editor” as used herein may include any user involved in the publication of the text content. An editor use case may include users who create and / or publish the text content, and intend for the text content to be available to end users in different consumption formats, without having to manually create the content in multiple formats. After an editor logs in via browser 14, web application 18 may then be opened, e.g. on the editor’s user device. Text content may then be selected by the editor, which may then be, uploaded, transformed and distributed as discussed in more detail below.
[0032] Editors and end users, e.g. consumers of content, may also access the format selection toolbar directly from a webpage, e.g. a page containing the text content to be presented to the end user in the one or more alternate formats. In some embodiments, embed toolbar 16 may launch the web application 18 for an editor, and may also launch the content, in the alternative format, for an end user. In some embodiments, embed toolbar 16 may present a link, such as a QR code, on a device other than mobile device 12, which can be captured by mobile device 12, e.g. via a camera, which will direct mobile device 12 to the alternative content.
[0033] Mobile device 12, web browser 14, and / or embed toolbar 16 may provide clients and users the ability to request some or all of the following transformations of the text content via a back-end API 20. Transformations may include Text Search Engine Optimization (SEO), Text Translations e.g. into other languages, text summary, audio generation, image generation, video generation, or American Sign Language generation. Other formats may also be used. Also, formatsPatent Docket No. 140276-5001may be used in combination. For example, a video version of the text content may also include a voice-over voice. Such a voice-over voice may also read a generated summary of the text content, rather than the original verbatim text content. Requests to back-end API 20 may be prioritized using an event queue 22.
[0034] In some embodiments, Web app 18 will send a request to the back-end API 20 to access one of Al models 24. In some embodiments, Al models 24 may be called upon directly. In other embodiments, Al models 24 may be called by agents, which communication may be by Application Program Interface (API). In some embodiments, prompts, as detailed below, may be sent to one or more agents, via an API. The agent may then pass the prompt to an Al model, which will return the result via the API. The request may include the original text content or a link or pointer thereto. The request may also include Al configuration parameters to instruct the Al model in how to create the transformed content, as described in further detail below.
[0035] Back-end API 20 may then send the text content to be converted, and may also send the Al configuration parameters, e.g. prompts, to the appropriate Al model 24 for the chosen content format. In some embodiments, some of the Al configuration parameters may be set by the editor, e.g. in accordance with a user interface that will be discussed in further detail below. For example, for Text SEO Optimization, Al configuration parameters may include Al prompt to generate a search engine optimized version of the source text. For text translations, Al configuration parameters may include Al prompt to generate text translations, and may include one or more language parameters. For text summary, Al configuration parameters may include an Al prompt to generate a text summary, which may include parameters relating to the length of the summary, e.g. in words. For audio generation: Al prompt to generate audio of the source text, and / or Al prompt to generate a text summary and then generate audio of the text summary. For audio of a text summary, parameters may include the length of the audio, in minutes, and a text summary may be created that is designed to be a length such that it takes the specified number of minutes to read the summary aloud.
[0036] For image generation, alternative approaches may be used. In one approach, an Al prompt to generate a static visual representation of the source text may be used. In another approach, an Al prompt to generate a script identifying key information to be shown in thePatent Docket No. 140276-5001visualization may be used. In this approach, API tools may be used to map information designated as key, to predefined visualization templates and generate the static visual representation.
[0037] For video generation, alternative approaches may also be used. In one approach, an Al prompt to generate a video representation of the source text may be used. In another approach, an Al prompt to generate a script identifying key information, and to identify stock footage to be captured in the video. API tools may be used to map the key information and stock footage to predefined video templates and generate the video representation. Video generation may include voice-over of either the full text of a summary. In the event of a summary, the parameters may include length parameters, similar to those discussed with respect to audio of a text summary, such that the voice-over voice reads a summary specifically constructed by the artificial intelligence engine to be read in the specified number of minutes.
[0038] Al configuration parameters for American Sign Language Generation may include using API tools to access Al generated American Sign Language representation of the text content and / or an Al generated summary of the text content.
[0039] After the alternative content is generated, back-end API 20 may receive the external content, and then may store the alternative content in a database 26 and / or file storage 28. A notification microservice 30 may also be used for approval workflow and requests for review and / or approval from back-end API 20. Review and approval information may be captured by application logs 32.
[0040] Upon package approval, web app 18 may create an embed code that generates or updates the embed toolbar 16. The embed code will be placed on the webpage of the source content. The embed toolbar 16 may then be updated with the embed code. In some embodiments, the embed code may be supplied as a JavaScript embed code or an iFrame embed code.
[0041] FIG. 2 illustrates a block diagram of a computing device 50, in accordance with some embodiments. In some embodiments, each of the mobile device 12 and browser 14, in FIG.1 may be, include, or use the features shown in FIG. 2. Servers, such as servers that run back-end API 20, notification microservice 30, artificial intelligence models 24, may also run on computing devices resembling computing device 50. Although FIG. 2 is described with respect to certainPatent Docket No. 140276-5001components shown therein, it will be appreciated that the elements of the computing device 50 may be combined, omitted, and / or replicated. In addition, it will be appreciated that additional elements other than those illustrated in FIG. 2 may be added to the computing device.
[0042] As shown in FIG. 2, the computing device 50 may include one or more processors 52, an instruction memory 54, a working memory 56, one or more input / output devices 58, a transceiver 60, one or more communication ports 62, a display 64 with a user interface 66, and an optional location device 68, all operatively coupled to one or more data buses 70. The data buses 70 allow for communication among the various components. The data buses 70 may include wired, or wireless, communication channels.
[0043] The one or more processors 52 may include any processing circuitry operable to control operations of the computing device 50. In some embodiments, the one or more processors 52 include one or more distinct processors, each having one or more cores (e.g., processing circuits). Each of the distinct processors may have the same or different structure. The one or more processors 52 may include one or more central processing units (CPUs), one or more graphics processing units (GPUs), application specific integrated circuits (ASICs), digital signal processors (DSPs), a chip multiprocessor (CMP), a network processor, an input / output (I / O) processor, a media access control (MAC) processor, a radio baseband processor, a co-processor, a microprocessor such as a complex instruction set computer (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, and / or a very long instruction word (VLIW) microprocessor, or other processing device. The one or more processors 52 may also be implemented by a controller, a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a programmable logic device (PLD), etc.
[0044] In some embodiments, the one or more processors 52 are configured to implement an operating system (OS) and / or various applications. Examples of an OS include, for example, operating systems generally known under various trade names such as Apple macOS™, Microsoft Windows™, Android™, Linux™, and / or any other proprietary or open-source OS. Examples of applications include, for example, network applications, local applications, data input / output applications, user interaction applications, etc.Patent Docket No. 140276-5001
[0045] The instruction memory 54 may store instructions that are accessed (e.g., read) and executed by at least one of the one or more processors 52. For example, the instruction memory 54 may be a non-transitory, computer-readable storage medium such as a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), flash memory (e.g. NOR and / or NAND flash memory), content addressable memory (CAM), polymer memory (e.g., ferroelectric polymer memory), phase-change memory (e.g., ovonic memory), ferroelectric memory, silicon-oxide-nitride-oxide-silicon (SONOS) memory, a removable disk, CD-ROM, any non-volatile memory, or any other suitable memory. The one or more processors 52 may be configured to perform a certain function or operation by executing code, stored on the instruction memory 54, embodying the function or operation. For example, the one or more processors 52 may be configured to execute code stored in the instruction memory 54 to perform one or more of any function, method, or operation disclosed herein.
[0046] Additionally, the one or more processors 52 may store data to, and read data from, the working memory 56. For example, the one or more processors 52 may store a working set of instructions to the working memory 56, such as instructions loaded from the instruction memory 54. The one or more processors 52 may also use the working memory 56 to store dynamic data created during one or more operations. The working memory 56 may include, for example, random access memory (RAM) such as a static random access memory (SRAM) or dynamic random access memory (DRAM), Double-Data-Rate DRAM (DDR -RAM), synchronous DRAM (SDRAM), an EEPROM, flash memory (e.g. NOR and / or NAND flash memory), content addressable memory (CAM), polymer memory (e.g., ferroelectric polymer memory), phasechange memory (e.g., ovonic memory), ferroelectric memory, silicon-oxide-nitride-oxi de-silicon (SONOS) memory, a removable disk, CD-ROM, any non-volatile memory, or any other suitable memory. Although embodiments are illustrated herein including separate instruction memory 54 and working memory 56, it will be appreciated that the computing device 50 may include a single memory unit configured to operate as both instruction memory and working memory. Further, although embodiments are discussed herein including non-volatile memory, it will be appreciated that computing device 50 may include volatile memory components in addition to at least one nonvolatile memory component.Patent Docket No. 140276-5001
[0047] In some embodiments, the instruction memory 54 and / or the working memory 56 includes an instruction set, in the form of a file for executing various methods, such as methods for a back-end API 20, an embed toolbar 16, a web app 18, or Al models 24, as described herein. The instruction set may be stored in any acceptable form of machine-readable instructions, including source code or various appropriate programming languages. Some examples of programming languages that may be used to store the instruction set include, but are not limited to: Java, JavaScript, C, C++, C#, Python, Objective-C, Visual Basic, .NET, HTML, CSS, SQL, NoSQL, Rust, Perl, etc. In some embodiments a compiler or interpreter is configured to convert the instruction set into machine executable code for execution by the one or more processors 52.
[0048] The input-output devices 58 may include any suitable device that allows for data input or output. For example, the input-output devices 58 may include one or more of a keyboard, a touchpad, a mouse, a stylus, a touchscreen, a physical button, a speaker, a microphone, a keypad, a click wheel, a motion sensor, a camera, and / or any other suitable input or output device.
[0049] The transceiver 60 and / or the communication port(s) 62 allow for communication with a network, e.g. a cellular network. In a cellular network, the transceiver 60 is configured to allow communications with the cellular network. In some embodiments, the transceiver 60 is selected based on the type of the communication network 22 the computing device 50 will be operating in. The one or more processors 52 are operable to receive data from, or send data to, a network, via the transceiver 60.
[0050] The communication port(s) 62 may include any suitable hardware, software, and / or combination of hardware and software that is capable of coupling the computing device 50 to one or more networks and / or additional devices. The communication port(s) 62 may be arranged to operate with any suitable technique for controlling information signals using a desired set of communications protocols, services, or operating procedures. The communication port(s) 62 may include the appropriate physical connectors to connect with a corresponding communications medium, whether wired or wireless, for example, a serial port such as a universal asynchronous receiver / transmitter (UART) connection, a Universal Serial Bus (USB) connection, or any other suitable communication port or connection. In some embodiments, the communication port(s) 62 allows for the programming of executable instructions in the instruction memory 54. In somePatent Docket No. 140276-5001embodiments, the communication port(s) 62 allow for the transfer (e.g., uploading or downloading) of data, such as machine learning model training data.
[0051] In some embodiments, the communication port(s) 62 are configured to couple the computing device 50 to a network. The network may include local area networks (LAN) as well as wide area networks (WAN) including without limitation Internet, wired channels, wireless channels, communication devices including telephones, computers, wire, radio, optical and / or other electromagnetic channels, and combinations thereof, including other devices and / or components capable of / associated with communicating data. For example, the communication environments may include in-body communications, various devices, and various modes of communications such as wireless communications, wired communications, and combinations of the same.
[0052] In some embodiments, the transceiver 60 and / or the communication port(s) 62 are configured to utilize one or more communication protocols. Examples of wired protocols may include, but are not limited to, Universal Serial Bus (USB) communication, RS-232, RS-422, RS-423, RS-485 serial protocols, FireWire, Ethernet, Fibre Channel, MIDI, ATA, Serial ATA, PCI Express, T-l (and variants), Industry Standard Architecture (ISA) parallel communication, Small Computer System Interface (SCSI) communication, or Peripheral Component Interconnect (PCI) communication, etc. Examples of wireless protocols may include, but are not limited to, the Institute of Electrical and Electronics Engineers (IEEE) 8O2.xx series of protocols, such as IEEE 802.11a / b / g / n / ac / ag / ax / be, IEEE 802.16, IEEE 802.20, GSM cellular radiotelephone system protocols with GPRS, CDMA cellular radiotelephone communication systems with IxRTT, EDGE systems, EV-DO systems, EV-DV systems, HSDPA systems, Wi-Fi Legacy, Wi-Fi 1 / 2 / 3 / 4 / 5 / 6 / 6E, wireless personal area network (PAN) protocols, Bluetooth Specification versions 5.0, 6, 7, legacy Bluetooth protocols, passive or active radio-frequency identification (RFID) protocols, Ultra-Wide Band (UWB), Digital Office (DO), Digital Home, Trusted Platform Module (TPM), ZigBee, etc.
[0053] The display 64 may be any suitable display, and may display the user interface 66. The user interfaces 66 may enable user interaction with browser 14, embed toolbar 16, or web app 18. For example, the user interface 66 may be a user interface for an application of a networkPatent Docket No. 140276-5001environment operator that allows a user to view and interact with the operator’s website. In some embodiments, a user may interact with the user interface 66 by engaging the input-output devices 58. In some embodiments, the display 64 may be a touchscreen, where the user interface 66 is displayed on the touchscreen.
[0054] The display 64 may include a screen such as, for example, a Liquid Crystal Display (LCD) screen, a light-emitting diode (LED) screen, an organic LED (OLED) screen, a movable display, a projection, etc. In some embodiments, the display 64 may include a coder / decoder, also known as Codecs, to convert digital media data into analog signals. For example, the visual peripheral output device may include video Codecs, audio Codecs, or any other suitable type of Codec.
[0055] The optional location device 68 may be communicatively coupled to a location network and operable to receive position data from the location network. For example, in some embodiments, the location device 68 includes a GPS device configured to receive position data identifying a latitude and longitude from one or more satellites of a GPS constellation. As another example, in some embodiments, the location device 68 is a cellular device configured to receive location data from one or more localized cellular towers. Based on the position data, the computing device 50 may determine a local geographical area (e.g., town, city, state, etc.) of its position.
[0056] In some embodiments, the computing device 50 is configured to implement one or more modules or engines, each of which is constructed, programmed, configured, or otherwise adapted, to autonomously carry out a function or set of functions. A module / engine may include a component or arrangement of components implemented using hardware, such as by an application specific integrated circuit (ASIC) or field-programmable gate array (FPGA), for example, or as a combination of hardware and software, such as by a microprocessor system and a set of program instructions that adapt the module / engine to implement the particular functionality, which (while being executed) transform the microprocessor system into a special-purpose device. A module / engine may also be implemented as a combination of the two, with certain functions facilitated by hardware alone, and other functions facilitated by a combination of hardware and software. In certain implementations, at least a portion, and in some cases, all, of a module / engine may be executed on the processor(s) of one or more computing platforms that are made up ofPatent Docket No. 140276-5001hardware e.g., one or more processors, data storage devices such as memory or drive storage, input / output facilities such as network interface devices, video devices, keyboard, mouse or touchscreen devices, etc.) that execute an operating system, system programs, and application programs, while also implementing the engine using multitasking, multithreading, distributed (e.g., cluster, peer-peer, cloud, etc.) processing where appropriate, or other such techniques. Accordingly, each module / engine may be realized in a variety of physically realizable configurations, and should generally not be limited to any particular implementation exemplified herein, unless such limitations are expressly called out. In addition, a module / engine may itself be composed of more than one sub- modules or sub-engines, each of which may be regarded as a module / engine in its own right. Moreover, in the embodiments described herein, each of the various modules / engines corresponds to a defined autonomous functionality; however, it should be understood that in other contemplated embodiments, each functionality may be distributed to more than one module / engine. Likewise, in other contemplated embodiments, multiple defined functionalities may be implemented by a single module / engine that performs those multiple functions, possibly alongside other functions, or distributed differently among a set of modules / engines than specifically illustrated in the embodiments herein.
[0057] Certain steps in a process for a transformation project, performed by an editor, in accordance with one aspect of the present disclosure, will now be described. An editor may begin by opening a project and uploading or entering, and confirming, text content to be transformed. The editor may then select the parameter inputs to be applied for each transformation Al model 24. After all parameters are set, the editor may then submit the project to return the “Transformation Package.” Parameters may then be mapped, e.g. by back-end API 20, to an appropriate Al model 24. “Text to Text”, “Text to Image”, “Text to Speech” and / or “Text to Video” transformations of the original text content may then be generated and returned to the view of the editor. The editor may then review the package and / or make comments, and send the transformations to reviewers, who may approve or make suggestions for changes, e.g. changes in Al configuration parameters.
[0058] Once a project is approved, an embed code for approved packages may be generated and distributed. The approved transformation project is stored including original text, transformed renditions and embed code. The editor can share the embed code to web editors forPatent Docket No. 140276-5001use in content page design, so that the alternative format content can be retrieved by the user, e g. via an alternative content user interface element on a web page, as described below.
[0059] Turning now to FIG. 3 A, a user interface 300 for creating alternative format content transformations is shown. In some embodiments, user interface 300 may be hosted and served by a server that runs web app 18 of FIG. 1. Persons having skill in the art will understand that the present disclosure is not limited to the specific interfaces disclosed herein or shown in FIGs. 3A through 3C. In some embodiments, interface 300 shows a screen for creating a new transformation project, after source content has been identified in a previous step in the process as shown by chooser 302. Chooser 304 allows the editor to choose parameters for text to text, text to speech, text to image and text to video transformation. A “text to text” transformation may give a consumer the option to read a shorter summary of the text based content, rather than the full version. Parameters for text to text may include a length of the summary text, which may be a word count.
[0060] Text to speech” transformation may give a consumer the option to listen to a spoken word version of the content. A spoken word version of the content may be a verbatim reading of the original content, or may be a spoken word reading of a summary of the content, which may also be generated by an artificial intelligence model. Parameters for “text to speech” may include various characteristics of the voice-over, such as a gender or other vocal characteristics of the voice, a speed at which the voice may speak (e.g. in word per minute), and a length, in minutes, of the spoken word transformation. If a length is specified, artificial intelligence models may be used to create a summary designed to take a specified number of minutes to read aloud at a specified vocal speed.
[0061] “Text to image” transformations may include creation of informational graphic, or “infographics,” that include shortened versions of the original text content, such as bullet points, charts, graphs, and / or the like. “Text to video” transformations may include creation of a video which includes graphics, stock video footage, or the like, as well as voice-over spoken word version of either the original text or a summary. In some embodiments, the summary and the voiceover spoken word may be created in accordance with parameters similar to those used for “text to speech” transformations.Patent Docket No. 140276-5001
[0062] Parameter drop downs 306 allow editors to choose parameters. As shown, parameters for text to video include video format, duration, voice-over voice characteristics, voiceover voice pitch, and voice speed. A video format parameter may include aspect ratio of the video, and may also include orientation information such as portrait, landscape, etc. A voice characteristics parameter may include a chosen gender of a voice, other voice characteristics such as vocal timbre, perceived age, etc. A voice pitch may be chosen as a parameter, e.g., high, medium, or low pitched voices. Voice speed parameters may be used to indicate how quickly the voice over voice speaks. Such parameters could be words per minute, or could be levels of speed such as very fast, fast, medium, slow, very slow.
[0063] Transformation thumbnails 308 show transformed content for each transformation type, after generation by the Al models, as described above, using the parameters selected in parameter drop downs 306. After the Editor reviews each transformation project type, the editor can click button 310 to send the transformation project to a reviewer for approval and publication. After the project has been approved, the interface appears as shown in FIG. 3B. Embed code user interface element 312 allows an embed code to be generated and copied, and then added to, e.g., the HTML code for the web page where the original text content can be published, so that a viewer of that content may select the alternative content format for consumption.
[0064] FIG. 3C shows a format selector user interface element, in the form of a toolbar 16, which may be embeddable in content web pages such as web page 314, and may be interactive. Toolbar 16 may include buttons to activate transmission and display or playback of video, text summary, audio, or informative graphic, as shown. As discussed above, content transformations may be stored at database 26 and / or fde storage 28 of FIG. 1, after having been created by one or more Al models 24, in some embodiments in accordance with editor-chosen parameters as shown in FIG. 3B. Interactions with toolbar 16, by a user, e.g. by clicking one of the buttons indicating a preferred format, may then result in the transmission of, and display of, the transformed content, e.g. on user device 12 or browser 14. In some embodiments, the alternative content may load as superimposed above the original web page containing the original text content. In other embodiments, the alternative content may load as a separate page, e.g. open instead of, or in addition to, the page containing the original text content.Patent Docket No. 140276-5001
[0065] Toolbar 16 may in some embodiments be embedded in the web page that shows the original pre-transformation text based content, so that users can interact with toolbar 16 to activate content in alternate formats.
[0066] FIGs. 4A and 4B show a flow chart of a method for creating and displaying alternative formats for text based content, in accordance with one aspect of the present disclosure. The method includes the steps of transmitting (402), to a format conversion server, e.g. a server running back-end API 20, a request to convert a unit of text based content to at least two consumption formats. At the format conversion server, the method includes transmitting (404) the unit of text based content, and a first artificial intelligence (“Al”) configuration parameter, to a first Al model, e.g. Al model 24, the first Al model being configured to convert the text based content to a first converted content in a first consumption format, transmitting (406) the unit text based content, and a second Al configuration parameter, to a second artificial intelligence model, the second Al model being configured to convert the text based content to a second converted content in a second consumption format, wherein the first consumption format and the second consumption format are each one of graphic, video, audio, and summarized text, and wherein the first consumption format is separate and distinct from the second consumption format, and storing (408) the substantially text based content, the first converted content, and the second converted content, e.g. in a database 26 and / or file storage 28.
[0067] The method further includes transmitting (410) data representing a user interface to a user device, the user interface including the text based content and a format selector interface element. The format selector interface may include a toolbar, which may contain interactive elements such as those displayed in toolbar 16 of FIG. 3C. The method further includes displaying (412) the text based content and the format selector interface element on the user device, which may be integrated, e.g. in the same web page. The method further includes, in response to a user interaction with a first portion of the format selector interface element, said first portion indicating the first consumption format, transmitting (414) data representing the first converted content to the user device. The method further includes displaying the first converted content on the user device. The method further includes, in response to a user interaction with a second portion of the format selector interface element, said second portion indicating the second consumption format, transmitting (416) data representing the second converted content to the user device. The methodPatent Docket No. 140276-5001further includes displaying the second converted content on the user device. For example, interacting with a “text to speech” button on toolbar 16 may result in transmission and display, or playback, of an audio fde containing a spoken word version of the original text content or of a summary of the content, which may be transmitted to the user device 12 or 14 after being retrieved from, e.g., fde storage 28. Interacting with a different button on toolbar 16 may result in an infographic being retrieved from fde storage 28, transmitted to user device 12 or 14, and displayed.
[0068] In some embodiments, the first Al configuration parameter and the second Al configuration parameter include editor preferences, received via user input prior to transmitting the request to the conversion server, the editor preferences for the first prompt are specific to the unit of text based content being transmitted and contain presentation parameters for the first consumption format; and the editor preferences for the second prompt are specific to the unit of text based content being transmitted and contain presentation parameters for the second consumption format. As discussed above, editor preferences may include length of an audio or video file, voice characteristics for spoken word or voice-over for video, and length of a text summary. Other parameters may also be used.
[0069] In some embodiments, the first consumption format is a video, and the presentation parameters are from the group consisting of video orientation, video duration, voice-over voice, voice-over pitch, and voice-over speed. In some embodiments, converting the unit of text based content into a first converted content includes creating a summary of the text based content, wherein a length of the summary is configured to match the video duration.
[0070] In some embodiments, the first consumption format is audio, and the presentation parameters are from the group consisting of duration, voice, voice pitch, and speed of speech. In some embodiments, converting the substantially text based content into a first converted content includes creating a summary of the unit of text based content, wherein the summary is configured to match the duration.
[0071] In some embodiments, the first consumption format is summarized text, wherein the presentation parameters include a text length, and wherein converting the substantially text based content into a first converted content includes creating a summary of the unit of text based content, wherein the summary is configured to match the length.Patent Docket No. 140276-5001
[0072] In some embodiments, the first consumption format is a graphic containing text, wherein the text contained in the graphic is a summary of the unit of text based content, generated by the first Al, and designed to fit on the graphic. In some embodiments, the graphic further contains at least one chart generated by the first Al model based on the unit of text based content.
[0073] In some embodiments, the first consumption format is summarized text, and wherein the first Al configuration parameter includes a specified length for the summarized text.
[0074] FIG. 5A illustrates an artificial neural network 500, utilized by at least some systems and methods described herein, in accordance with some embodiments. Alternative terms for “artificial neural network” are “neural network,” “artificial neural net,” “neural net,” or “trained function.” The neural network 500 comprises nodes 520-544 and edges 546-548, wherein each edge 546-548 is a directed connection from a first node 520-538 to a second node 532-544. In general, the first node 520-538 and the second node 532-544 are different nodes, although it is also possible that the first node 520-538 and the second node 532-544 are identical. For example, in FIG. 3 the edge 546 is a directed connection from the node 520 to the node 532, and the edge 548 is a directed connection from the node 532 to the node 540. An edge 546-548 from a first node 520-538 to a second node 532-544 is also denoted as “ingoing edge” for the second node 532-544 and as “outgoing edge” for the first node 520-538.
[0075] The nodes 520-544 of the neural network 500 may be arranged in layers 510-514, wherein the layers may comprise an intrinsic order introduced by the edges 546-548 between the nodes 520-544 such that edges 546-548 exist only between neighboring layers of nodes. In the illustrated embodiment, there is an input layer 510 comprising only nodes 520-530 without an incoming edge, an output layer 514 comprising only nodes 540-544 without outgoing edges, and a hidden layer 512 in-between the input layer 510 and the output layer 514. In general, the number of hidden layer 512 may be chosen arbitrarily and / or through training. The number of nodes 520-530 within the input layer 510 usually relates to the number of input values of the neural network, and the number of nodes 540-544 within the output layer 514 usually relates to the number of output values of the neural network.
[0076] In particular, a (real) number may be assigned as a value to every node 520-544 of the neural network 500. Here, x^7denotes the value of the i-th node 520-544 of the n-th layer 510-Patent Docket No. 140276-5001514. The values of the nodes 520-530 of the input layer 510 are equivalent to the input values of the neural network 500, the values of the nodes 540-544 of the output layer 514 are equivalent to the output value of the neural network 500. Furthermore, each edge 546-548 may comprise a weight being a real number, in particular, the weight is a real number within the interval [-1, 1], within the interval [0, 1], and / or within any other suitable interval. Here, wi^’n'>denotes the weight of the edge between the i-th node 520-538 of the m-th layer 510, 512 and the j-th node 532-544 of the n-th layer 512, 514. Furthermore, the abbreviation w> is defined for the weight w <^->J’n+1
[0077] In particular, to calculate the output values of the neural network 500, the input values are propagated through the neural network. In particular, the values of the nodes 532-544 of the (n+l)-th layer 512, 514 may be calculated based on the values of the nodes 520-538 of the n-th layer 510, 512 byHerein, the function f is a transfer function (another term is “activation function”). Known transfer functions are step functions, sigmoid function (e.g., the logistic function, the generalized logistic function, the hyperbolic tangent, the Arctangent function, the error function, the smooth step function) or rectifier functions. The transfer function is mainly used for normalization purposes.
[0078] In particular, the values are propagated lay er- wise through the neural network, wherein values of the input layer 510 are given by the input of the neural network 500, wherein values of the hidden layer(s) 512 may be calculated based on the values of the input layer 510 of the neural network and / or based on the values of a prior hidden layer, etc.
[0079] In order to set the values‘-J for the edges, the neural network 500 has to be trained using training data. In particular, training data comprises training input data and training output data. For a training step, the neural network 500 is applied to the training input data to generate calculated output data. In particular, the training data and the calculated output data comprise a number of values, said number being equal with the number of nodes of the output layer.Patent Docket No. 140276-5001
[0080] In particular, a comparison between the calculated output data and the training data is used to recursively adapt the weights within the neural network 500 (backpropagation algorithm). In particular, the weights are changed according to> >( ?iwherein y is a learning rate, and the numbersmay be recursively calculated as< based on bn+1if the (n+l)-th layer is not the output layer, and> "if the (n+l)-th layer is the output layer 514, wherein f is the first derivative of the activation function, a is the comparison training value for the j-th node of the output layer 514.
[0081] FIG. 5B illustrates a tree-based neural network 550, utilized by systems and methods described herein, in accordance with some embodiments. In particular, the tree-based neural network 550 is a random forest neural network, though it will be appreciated that the discussion herein is applicable to other decision tree neural networks. The tree-based neural network 550 includes a plurality of trained decision trees 554a-554c each including a set of nodes 556 (also referred to as “leaves”) and a set of edges 558 (also referred to as “branches”).
[0082] Each of the trained decision trees 554a-554c may include a classification and / or a regression tree (CART). Classification trees include a tree model in which a target variable may take a discrete set of values, e.g., may be classified as one of a set of values. In classification trees, each leaf 556 represents class labels and each of the branches 558 represents conjunctions of features that connect the class labels. Regression trees include a tree model in which the target variable may take continuous values (e.g., a real number value).Patent Docket No. 140276-5001
[0083] In operation, an input data set 552 including one or more features or attributes is received. A subset of the input data set 552 is provided to each of the trained decision trees 554a-554c. The subset may include a portion of and / or all of the features or attributes included in the input data set 552. Each of the trained decision trees 554a-554c is trained to receive the subset of the input data set 552 and generate a tree output value 560a-560c, such as a classification or regression output. The individual tree output value 560a-560c is determined by traversing the trained decision trees 554a-554c to arrive at a final leaf (or node) 556.
[0084] In some embodiments, the tree-based neural network 550 applies an aggregation process 562 to combine the output of each of the trained decision trees 554a-554c into a final output 564. For example, in embodiments including classification trees, the tree-based neural network 550 may apply a maj ority -voting process to identify a classification selected by the majority of the trained decision trees 554a-554c. As another example, in embodiments including regression trees, the tree-based neural network 550 may apply an average, mean, and / or other mathematical process to generate a composite output of the trained decision trees. The final output 564 is provided as an output of the tree-based neural network 550.
[0085] FIG. 5C illustrates a deep neural network (DNN) 570, utilized by systems and methods described herein, in accordance with some embodiments. The DNN 570 is an artificial neural network, such as the neural network 500 illustrated in conjunction with FIG. 3, that includes representation learning. The DNN 570 may include an unbounded number of (e.g., two or more) intermediate layers 574a-574d each of a bounded size (e.g., having a predetermined number of nodes), providing for practical application and optimized implementation of a universal classifier. Each of the layers 574a-574d may be heterogenous. The DNN 570 may be configured to model complex, non-linear relationships. Intermediate layers, such as intermediate layer 574c, may provide compositions of features from lower layers, such as layers 574a, 574b, providing for modeling of complex data.
[0086] In some embodiments, the DNN 570 may be considered a stacked neural network including multiple layers each configured to execute one or more computations. The computation for a network with L hidden layers may be denoted as:Patent Docket No. 140276-5001where a^ x) is a preactivation function and h^(x) is a hidden-layer activation function providing the output of each hidden layer. The preactivation function a^{x) may include a linear operation with matrixand bias b^, where:
[0087] In some embodiments, the DNN 570 is a feedforward network in which data flows from an input layer 572 to an output layer 576 without looping back through any layers. In some embodiments, the DNN 570 may include a backpropagation network in which the output of at least one hidden layer is provided, e.g, propagated, to a prior hidden layer. The DNN 570 may include any suitable neural network, such as a self-organizing neural network, a recurrent neural network, a convolutional neural network, a modular neural network, and / or any other suitable neural network.
[0088] In some embodiments, a DNN 570 may include a neural additive model (NAM). An NAM includes a linear combination of networks, each of which attends to (e.g., provides a calculation regarding) a single input feature. For example, a NAM may be represented as:y = + (M) + fe) + ••• + Afe)where ft is an offset and each ft is parametrized by a neural network. In some embodiments, the DNN 570 may include a neural multiplicative model (NMM), including a multiplicative form for the NAM mode using a log transformation of the dependent variable y and the independent variable x:where d represents one or more features of the independent variable x.Patent Docket No. 140276-5001
[0089] In some embodiments, a Al models 24 can include and / or implement one or more trained models, such as a text to text, text to speech, text to video, or text to graphic models as discussed herein. In some embodiments, one or more trained models can be generated using an iterative training process based on a training dataset. FIG. 6A illustrates a method 600 for generating a trained model, such as a trained text to speech, text to text, text to video, or text to graphic model, in accordance with some embodiments. FIG. 6B is a process flow 650 illustrating various steps of the method 600 of generating a trained model, in accordance with some embodiments. At step 602, a training dataset 652 is received by a system, such as a processing device 10. The training dataset 652 can include labeled and / or unlabeled data. For example, in some embodiments, a set of original text content may be provided for use in training a model.
[0090] At optional step 604, the received training dataset 652 is processed and / or normalized by a normalization module 660. For example, in some embodiments, the training dataset 652 can be augmented by imputing or estimating missing values of one or more features associated with format transformation models.
[0091] At step 606, an iterative training process is executed to train a selected model framework 662. The selected model framework 662 can include an untrained (e.g., base) machine learning model and / or a partially or previously trained model (e.g., a prior version of a trained model). The training process is configured to iteratively adjust parameters (e.g., hyperparameters) of the selected model framework 662 to minimize a cost value (e.g., an output of a cost function) for the selected model framework 662.
[0092] The training process is an iterative process that generates set of revised model parameters 666 during each iteration. The set of revised model parameters 666 can be generated by applying an optimization process 664 to the cost function of the selected model framework 662. The optimization process 664 can be configured to reduce the cost value (e.g., reduce the output of the cost function) at each step by adjusting one or more parameters during each iteration of the training process.
[0093] After each iteration of the training process, at step 608, a determination is made whether the training process is complete. The determination at step 608 can be based on any suitable parameters. For example, in some embodiments, a training process can complete after aPatent Docket No. 140276-5001predetermined number of iterations. As another example, in some embodiments, a training process can complete when it is determined that the cost function of the selected model framework 662 has reached a minimum, such as a local minimum and / or a global minimum.
[0094] At step 610, a trained model 668, such as a trained format transformation, is output and provided for use in a method of transforming text, such as the method 600 discussed above with respect to FIGS. 6-7. At optional step 612, a trained model 668 can be evaluated by an evaluation process 670. A trained model can be evaluated based on any suitable metrics, such as, for example, an F or F 1 score, normalized discounted cumulative gain (NDCG) of the model, mean reciprocal rank (MRR), mean average precision (MAP) score of the model, and / or any other suitable evaluation metrics. Although specific embodiments are discussed herein, it will be appreciated that any suitable set of evaluation metrics can be used to evaluate a trained model.
[0095] FIG. 7 is a structural diagram of an LLM 700 formed in a transformer architecture, utilized by systems and methods described herein, in accordance with some embodiments. Generative Al uses natural language processing (NLP) and machine learning to create natural language data or content. The LLM 700 includes a deep learning neural network configured to perform natural language processing (NLP) tasks, e.g., text generation, summarization, translation, text classification, and answering questions, thereby enabling Generative Al. Compared with a normal language model, the LLM 700 includes more than 100 million parameters and is pretrained with on large corpora of text. In some embodiments, the LLM 700 is implemented with a transformer architecture, and configured to shift through large datasets and recognize patterns and relationships between words or phrases. The transformer architecture includes an attention mechanism that weighs the importance of different words or phrases in a given context.
[0096] The transformer architecture of the LLM 700 includes an encoder network 702 and a decoder network 704. The encoder network 702 is configured to receive an input sequence 706 and generate a sequence of hidden states 708. Each hidden state 708 includes a vector that encodes contextual information of a word in the input sequence 706 based on their relative positions. The decoder network 704 is configured to receive portions 710P of a target sequence 710 successively and use an output 712 of the encoder network 702 to generate the target sequence 710. In some embodiments, the decoder network 704 starts with a starting token (e.g., “start”) and generates onePatent Docket No. 140276-5001prediction at a time. The decoder network 704 uses the output 712 produced by the encoder network 702 to understand the context of the input sequence 706. For each word of the target sequence 710 to be predicted, the decoder network 704 uses cross-attention mechanisms to focus on corresponding portions of the output 712 of the encoder network 702. As each word of the target sequence 710 is generated, the decoder network 704 updates its state and predicts a next word, until the entire target sequence 710 is generated.
[0097] In some embodiments, the LLM 700 applies a self-attention mechanism, and each position in a sequence (e.g., a natural language query) is attended to all positions in the same sequence. Self-attention helps the LLM 700 to understand and interpret the sequence by considering the entire sequence. For instance, when processing the natural language query, selfattention allows each word to be contextualized in relation to every other word in that natural language query. Alternatively, in some embodiments, the LLM 700 applies a transformer architecture including multihead attention 720 (also called multihead self-attention). Each attention head 722 learns a respective attention mechanism so that multihead attention 720 as a whole can learn more complex relationships. For example, referring to Figure 3, multihead attention 720 is applied in both the encoder network 702 and the decoder network 704 of the LLM 700 implemented in the transformer architecture.
[0098] Although the subject matter has been described in terms of exemplary embodiments, it is not limited thereto. Rather, the appended claims should be construed broadly, to include other variants and embodiments, which may be made by those skilled in the art.
[0099] The following describes some example implementations.
[0100] In some implementations, a system of one or more computers can be configured to perform particular operations or actions by virtue of having software, firmware, hardware, or a combination of them installed on the system that in operation causes or cause the system to perform the actions. One or more computer programs can be configured to perform particular operations or actions by virtue of including instructions that, when executed by data processing apparatus, cause the apparatus to perform the actions. One general aspect includes a computer-implemented method. The computer-implemented method also includes transmitting, to a format conversion server, a request to convert a unit of text based content to at least two consumption formats. ThePatent Docket No. 140276-5001method also includes at the format conversion server:. The method also includes transmitting the unit of text based content, and a first artificial intelligence (ai) configuration parameter, to a first artificial intelligence model, the first ai model being configured to convert the text based content to a first converted content in a first consumption format. The method also includes transmitting the unit text based content, and a second ai configuration parameter, to a second artificial intelligence model, the second ai model being configured to convert the text based content to a second converted content in a second consumption format, where the first consumption format and the second consumption format are each one of graphic, video, audio, and summarized text, and where the first consumption format is separate and distinct from the second consumption format. The method also includes storing the substantially text based content, the first converted content, and the second converted content. The method also includes transmitting data representing a user interface to a user device, the user interface including the text based content and a format selector interface element. The method also includes displaying the text based content and the format selector interface element on the user device. The method also includes in response to a user interaction with a first portion of the format selector interface element, said first portion indicating the first consumption format, transmitting data representing the first converted content to the user device, and displaying the first converted content on the user device. The method also includes in response to a user interaction with a second portion of the format selector interface element, said second portion indicating the second consumption format, transmitting data representing the second converted content to the user device, and displaying the second converted content on the user device. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods.
[0101] Some implementations of the computer-implemented method may include one or more of the following features. The method where the first ai configuration parameter and the second ai configuration parameter include editor preferences, received via user input prior to transmitting the request to the conversion server, the editor preferences for the first prompt are specific to the unit of text based content being transmitted and contain presentation parameters for the first consumption format; and the editor preferences for the second prompt are specific to the unit of text based content being transmitted and contain presentation parameters for the second consumption format. The first consumption format is a video, and where the presentationPatent Docket No. 140276-5001parameters are from the group may include of video orientation, video duration, voice-over voice, voice-over pitch, and voice-over speed. Converting the unit of text based content into a first converted content includes creating a summary of the text based content, where a length of the summary is configured to match the video duration. The first consumption format is audio, and where the presentation parameters are from the group may include of duration, voice, voice pitch, and speed of speech. Converting the substantially text based content into a first converted content includes creating a summary of the unit of text based content, where the summary is configured to match the duration. The first consumption format is summarized text, where the presentation parameters include a text length, and where converting the substantially text based content into a first converted content includes creating a summary of the unit of text based content, where the summary is configured to match the length. The first consumption format is a graphic containing text, where the text contained in the graphic is a summary of the unit of text based content, generated by the first ai, and designed to fit on the graphic. The graphic further contains at least one chart generated by the first ai model based on the unit of text based content. The first consumption format is summarized text, and where the first ai configuration parameter includes a specified length for the summarized text. Implementations of the described techniques may include hardware, a method or process, or computer software on a computer-accessible medium.
[0102] Some implementation include a system. The system also includes a non-transitory memory. The system also includes a processor communicatively coupled to the non-transitory memory, where the processor is configured to read a set of instructions to:. The system also includes transmit, to a format conversion server, a request to convert a unit of text based content to at least two consumption formats. The system also includes at the format conversion server:. The system also includes transmit the unit of text based content, and a first artificial intelligence (ai) configuration parameter, to a first artificial intelligence model, the first ai model being configured to convert the text based content to a first converted content in a first consumption format. The system also includes transmit the unit text based content, and a second ai configuration parameter, to a second artificial intelligence model, the second ai model being configured to convert the text based content to a second converted content in a second consumption format, where the first consumption format and the second consumption format are each one of graphic, video, audio, and summarized text, and where the first consumption format is separate and distinctPatent Docket No. 140276-5001from the second consumption format. The system also includes store the substantially text based content, the first converted content, and the second converted content. The system also includes transmit data representing a user interface to a user device, the user interface including the text based content and a format selector interface element. The system also includes display the text based content and the format selector interface element on the user device. The system also includes in response to a user interaction with a first portion of the format selector interface element, said first portion indicating the first consumption format, transmit data representing the first converted content to the user device, and displaying the first converted content on the user device. The system also includes in response to a user interaction with a second portion of the format selector interface element, said second portion indicating the second consumption format, transmit data representing the second converted content to the user device, and displaying the second converted content on the user device. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods.
[0103] Some implementations of the system may include one or more of the following features. The system where the first ai configuration parameter and the second ai configuration parameter include editor preferences, received via user input prior to transmitting the request to the conversion server, the editor preferences for the first prompt are specific to the unit of text based content being transmitted and contain presentation parameters for the first consumption format; and the editor preferences for the second prompt are specific to the unit of text based content being transmitted and contain presentation parameters for the second consumption format. The first consumption format is a video, and where the presentation parameters are from the group may include of video orientation, video duration, voice-over voice, voice-over pitch, and voiceover speed. Converting the unit of text based content into a first converted content includes creating a summary of the text based content, where a length of the summary is configured to match the video duration. The first consumption format is audio, and where the presentation parameters are from the group may include of duration, voice, voice pitch, and speed of speech. Converting the substantially text based content into a first converted content includes creating a summary of the unit of text based content, where the summary is configured to match the duration. The first consumption format is summarized text, where the presentation parameters include a text length, and where converting the substantially text based content into a first converted content includesPatent Docket No. 140276-5001creating a summary of the unit of text based content, where the summary is configured to match the length. The first consumption format is a graphic containing text, where the text contained in the graphic is a summary of the unit of text based content, generated by the first ai, and designed to fit on the graphic. The graphic further contains at least one chart generated by the first ai model based on the unit of text based content. The first consumption format is summarized text, and where the first ai configuration parameter includes a specified length for the summarized text.
[0104] Some implementation include a non-transitory computer readable medium having instructions stored thereon. The non-transitory computer readable medium also includes transmitting, to a format conversion server, a request to convert a unit of text based content to at least two consumption formats. The medium also includes at the format conversion server:. The medium also includes transmitting the unit of text based content, and a first artificial intelligence (ai) configuration parameter, to a first artificial intelligence model, the first ai model being configured to convert the text based content to a first converted content in a first consumption format. The medium also includes transmitting the unit text based content, and a second ai configuration parameter, to a second artificial intelligence model, the second ai model being configured to convert the text based content to a second converted content in a second consumption format, where the first consumption format and the second consumption format are each one of graphic, video, audio, and summarized text, and where the first consumption format is separate and distinct from the second consumption format. The medium also includes storing the substantially text based content, the first converted content, and the second converted content. The medium also includes transmitting data representing a user interface to a user device, the user interface including the text based content and a format selector interface element. The medium also includes displaying the text based content and the format selector interface element on the user device. The medium also includes in response to a user interaction with a first portion of the format selector interface element, said first portion indicating the first consumption format, transmitting data representing the first converted content to the user device, and displaying the first converted content on the user device. The medium also includes in response to a user interaction with a second portion of the format selector interface element, said second portion indicating the second consumption format, transmitting data representing the second converted content to the user device, and displaying the second converted content on the user device. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programsPatent Docket No. 140276-5001recorded on one or more computer storage devices, each configured to perform the actions of the methods.
[0105] Some implementations of the non-transitoiy computer readable medium may include one or more of the following features. The method where the first ai configuration parameter and the second ai configuration parameter include editor preferences, received via user input prior to transmitting the request to the conversion server, the editor preferences for the first prompt are specific to the unit of text based content being transmitted and contain presentation parameters for the first consumption format; and the editor preferences for the second prompt are specific to the unit of text based content being transmitted and contain presentation parameters for the second consumption format. The first consumption format is a video, and where the presentation parameters are from the group may include of video orientation, video duration, voice-over voice, voice-over pitch, and voice-over speed. Converting the unit of text based content into a first converted content includes creating a summary of the text based content, where a length of the summary is configured to match the video duration. The first consumption format is audio, and where the presentation parameters are from the group may include of duration, voice, voice pitch, and speed of speech. Converting the substantially text based content into a first converted content includes creating a summary of the unit of text based content, where the summary is configured to match the duration. The first consumption format is summarized text, where the presentation parameters include a text length, and where converting the substantially text based content into a first converted content includes creating a summary of the unit of text based content, where the summary is configured to match the length. The first consumption format is a graphic containing text, where the text contained in the graphic is a summary of the unit of text based content, generated by the first ai, and designed to fit on the graphic. The graphic further contains at least one chart generated by the first ai model based on the unit of text based content. The first consumption format is summarized text, and where the first ai configuration parameter includes a specified length for the summarized text. Implementations of the described techniques may include hardware, a method or process, or computer software on a computer-accessible medium.
Claims
Patent Docket No. 140276-5001 ClaimsWhat is claimed is:
1. A computer-implemented method, comprising:transmitting, to a format conversion server, a request to convert a unit of text based content to at least two consumption formats;at the format conversion server:transmitting the unit of text based content, and a first artificial intelligence (“Al”) configuration parameter, to a first artificial intelligence model, the first Al model being configured to convert the text based content to a first converted content in a first consumption format;transmitting the unit text based content, and a second Al configuration parameter, to a second artificial intelligence model, the second Al model being configured to convert the text based content to a second converted content in a second consumption format, wherein the first consumption format and the second consumption format are each one of graphic, video, audio, and summarized text, and wherein the first consumption format is separate and distinct from the second consumption format;storing the substantially text based content, the first converted content, and the second converted content;transmitting data representing a user interface to a user device, the user interface including the text based content and a format selector interface element;displaying the text based content and the format selector interface element on the user device; in response to a user interaction with a first portion of the format selector interface element, said first portion indicating the first consumption format, transmitting data representing the first converted content to the user device, and displaying the first converted content on the user device; and in response to a user interaction with a second portion of the format selector interface element, said second portion indicating the second consumption format, transmitting data representing the second converted content to the user device, and displaying the second converted content on the user device.
2. The method of claim 1, whereinthe first Al configuration parameter and the second Al configuration parameter include editor preferences, received via user input prior to transmitting the request to the conversion server,Patent Docket No. 140276-5001 the editor preferences for the first prompt are specific to the unit of text based content being transmitted and contain presentation parameters for the first consumption format; andthe editor preferences for the second prompt are specific to the unit of text based content being transmitted and contain presentation parameters for the second consumption format.
3. The method of claim 2, wherein the first consumption format is a video, and wherein the presentation parameters are from the group consisting of video orientation, video duration, voice-over voice, voice-over pitch, and voice-over speed.
4. The method of claim 3, wherein converting the unit of text based content into a first converted content includes creating a summary of the text based content, wherein a length of the summary is configured to match the video duration.
5. The method of claim 2, wherein the first consumption format is audio, and wherein the presentation parameters are from the group consisting of duration, voice, voice pitch, and speed of speech.
6. The method of claim 5, wherein converting the substantially text based content into a first converted content includes creating a summary of the unit of text based content, wherein the summary is configured to match the duration.
7. The method of claim 2, wherein the first consumption format is summarized text, wherein the presentation parameters include a text length, and wherein converting the substantially text based content into a first converted content includes creating a summary of the unit of text based content, wherein the summary is configured to match the length.
8. The method of claim 2, wherein the first consumption format is a graphic containing text, wherein the text contained in the graphic is a summary of the unit of text based content, generated by the first Al, and designed to fit on the graphic.
9. The method of claim 8, wherein the graphic further contains at least one chart generated by the first Al model based on the unit of text based content.Patent Docket No. 140276-5001 10. The method of claim 2, wherein the first consumption format is summarized text, and wherein the first Al configuration parameter includes a specified length for the summarized text.
11. A system, comprising:a non-transitory memory;a processor communicatively coupled to the non-transitory memory, wherein the processor is configured to read a set of instructions to:transmit, to a format conversion server, a request to convert a unit of text based content to at least two consumption formats;at the format conversion server:transmit the unit of text based content, and a first artificial intelligence (“Al”) configuration parameter, to a first artificial intelligence model, the first Al model being configured to convert the text based content to a first converted content in a first consumption format;transmit the unit text based content, and a second Al configuration parameter, to a second artificial intelligence model, the second Al model being configured to convert the text based content to a second converted content in a second consumption format, wherein the first consumption format and the second consumption format are each one of graphic, video, audio, and summarized text, and wherein the first consumption format is separate and distinct from the second consumption format;store the substantially text based content, the first converted content, and the second converted content;transmit data representing a user interface to a user device, the user interface including the text based content and a format selector interface element;display the text based content and the format selector interface element on the user device;in response to a user interaction with a first portion of the format selector interface element, said first portion indicating the first consumption format, transmit data representing the first converted content to the user device, and displaying the first converted content on the user device; and in response to a user interaction with a second portion of the format selector interface element, said second portion indicating the second consumption format, transmit data representing the second converted content to the user device, and displaying the second converted content on the user device.Patent Docket No. 140276-5001 12. The system of claim 11, whereinthe first Al configuration parameter and the second Al configuration parameter include editor preferences, received via user input prior to transmitting the request to the conversion server, the editor preferences for the first prompt are specific to the unit of text based content being transmitted and contain presentation parameters for the first consumption format; andthe editor preferences for the second prompt are specific to the unit of text based content being transmitted and contain presentation parameters for the second consumption format.
13. The system of claim 12, wherein the first consumption format is a video, and wherein the presentation parameters are from the group consisting of video orientation, video duration, voice-over voice, voice-over pitch, and voice-over speed.
14. The system of claim 13, wherein converting the unit of text based content into a first converted content includes creating a summary of the text based content, wherein a length of the summary is configured to match the video duration.
15. The system of claim 12, wherein the first consumption format is audio, and wherein the presentation parameters are from the group consisting of duration, voice, voice pitch, and speed of speech.
16. The system of claim 15, wherein converting the substantially text based content into a first converted content includes creating a summary of the unit of text based content, wherein the summary is configured to match the duration.
17. The system of claim 12, wherein the first consumption format is summarized text, wherein the presentation parameters include a text length, and wherein converting the substantially text based content into a first converted content includes creating a summary of the unit of text based content, wherein the summary is configured to match the length.
18. The system of claim 12, wherein the first consumption format is a graphic containing text, wherein the text contained in the graphic is a summary of the unit of text based content, generated by the first Al, and designed to fit on the graphic.Patent Docket No. 140276-5001 19. The system of claim 18, wherein the graphic further contains at least one chart generated by the first Al model based on the unit of text based content.
20. The system of claim 12, wherein the first consumption format is summarized text, and wherein the first Al configuration parameter includes a specified length for the summarized text.
21. A non-transitory computer readable medium having instructions stored thereon, wherein the instructions, when executed by at least one processor, cause at least one device to perform operations comprising:transmitting, to a format conversion server, a request to convert a unit of text based content to at least two consumption formats;at the format conversion server:transmitting the unit of text based content, and a first artificial intelligence (“Al”) configuration parameter, to a first artificial intelligence model, the first Al model being configured to convert the text based content to a first converted content in a first consumption format;transmitting the unit text based content, and a second Al configuration parameter, to a second artificial intelligence model, the second Al model being configured to convert the text based content to a second converted content in a second consumption format, wherein the first consumption format and the second consumption format are each one of graphic, video, audio, and summarized text, and wherein the first consumption format is separate and distinct from the second consumption format;storing the substantially text based content, the first converted content, and the second converted content;transmitting data representing a user interface to a user device, the user interface including the text based content and a format selector interface element;displaying the text based content and the format selector interface element on the user device; in response to a user interaction with a first portion of the format selector interface element, said first portion indicating the first consumption format, transmitting data representing the first converted content to the user device, and displaying the first converted content on the user device; and in response to a user interaction with a second portion of the format selector interface element, said second portion indicating the second consumption format, transmitting data representing the second converted content to the user device, and displaying the second converted content on the user device.Patent Docket No. 140276-500122. The method of claim 21, whereinthe first Al configuration parameter and the second Al configuration parameter include editor preferences, received via user input prior to transmitting the request to the conversion server, the editor preferences for the first prompt are specific to the unit of text based content being transmitted and contain presentation parameters for the first consumption format; andthe editor preferences for the second prompt are specific to the unit of text based content being transmitted and contain presentation parameters for the second consumption format.
23. The medium of claim 22, wherein the first consumption format is a video, and wherein the presentation parameters are from the group consisting of video orientation, video duration, voice-over voice, voice-over pitch, and voice-over speed.
24. The medium of claim 23, wherein converting the unit of text based content into a first converted content includes creating a summary of the text based content, wherein a length of the summary is configured to match the video duration.
25. The medium of claim 22, wherein the first consumption format is audio, and wherein the presentation parameters are from the group consisting of duration, voice, voice pitch, and speed of speech.
26. The medium of claim 25, wherein converting the substantially text based content into a first converted content includes creating a summary of the unit of text based content, wherein the summary is configured to match the duration.
27. The medium of claim 22, wherein the first consumption format is summarized text, wherein the presentation parameters include a text length, and wherein converting the substantially text based content into a first converted content includes creating a summary of the unit of text based content, wherein the summary is configured to match the length.
28. The medium of claim 22, wherein the first consumption format is a graphic containing text, wherein the text contained in the graphic is a summary of the unit of text based content, generated by the first Al, and designed to fit on the graphic.Patent Docket No. 140276-500129. The medium of claim 28, wherein the graphic further contains at least one chart generated by the first Al model based on the unit of text based content.
30. The medium of claim 22, wherein the first consumption format is summarized text, and wherein the first Al configuration parameter includes a specified length for the summarized text.