system
The system uses natural language processing to generate diverse digital content in various styles and genres, addressing the limitations of conventional methods by providing multilingual support and collaboration features, thus enhancing efficiency and creativity.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-21
- Publication Date
- 2026-05-07
AI Technical Summary
Conventional methods for generating digital content are time-consuming, require specialized skills, and lack multilingual support, genre and style flexibility, and efficient collaboration features.
A system utilizing natural language processing to analyze text and generate visual and auditory content in multiple styles and genres, with multilingual support and collaboration tools, allowing easy sharing and editing among users.
Significantly reduces time and effort in content creation, enhances creativity, and facilitates efficient multilingual and collaborative content generation.
Smart Images

Figure 2026074999000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In modern digital content production, it is required to generate visual and auditory elements with high quality and efficiency. However, conventional methods require a lot of time and specialized skills, and it is particularly difficult to provide options for multilingual support, different styles, and genres. In addition, since it is not easy to easily share or co-edit the created content among users, there are limitations in improving the efficiency and promoting the creativity of content production.
Means for Solving the Problems
[0005] This invention provides a system that uses natural language processing technology to analyze the meaning of text and automatically generates visual and auditory digital content in various styles and genres based on the analysis results. This enables multilingual support and allows content generation according to the language and style selected by the user. Furthermore, it incorporates collaboration functions, allowing users to easily share and collaboratively edit the content they have generated. This significantly reduces time and effort and promotes the creative process.
[0006] "Natural language processing technology" is a technology that enables computers to understand, interpret, and manipulate human language.
[0007] "Visual digital content" refers to content such as images, illustrations, and videos that are presented visually in digital format.
[0008] "Auditory digital content" refers to content such as audio, music, and narration that is presented audibly in digital format.
[0009] "Multilingual support" refers to a specification or function that allows operation and display in multiple languages.
[0010] "Style" refers to a specific aesthetic expression or design trend in content creation.
[0011] "Genre" refers to a criterion used to classify themes and categories within a work or piece of content.
[0012] "Collaboration tools" are features that enable multiple users to collaboratively create, edit, and share content.
[0013] A "user" refers to a person or group that uses the system to generate or manipulate content.
[0014] "Sharing" refers to the act of enabling the use or display of the generated content together with other users.
[0015] "Collaborative editing" refers to multiple users simultaneously editing or modifying the same content.
Brief Description of Drawings
[0016] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which multiple emotions are mapped. [Figure 10] It shows an emotion map to which multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13]It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.
Modes for Carrying Out the Invention
[0017] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be one arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be one type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0020] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0021] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0024] [First Embodiment]
[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0037] This invention provides users with a system for generating visual and auditory digital content using natural language processing technology. The following is an example of this system.
[0038] User interface and text input
[0039] The user accesses the system using their device and logs in to begin the content generation process. A text input form appears on the device, and the user enters text related to the content they want to generate into this form. For example, this could be a portion of a novel or the beginning of a blog post.
[0040] Data reception and analysis
[0041] The server receives text data sent from the user's terminal. The received text is parsed by the server's natural language processing (NLP) module. Specifically, it performs tokenization, syntactic analysis, and semantic analysis of the text to determine what visual and auditory content is appropriate based on its characteristics and context.
[0042] Style and Genre Selection
[0043] Users can also select the style and genre of the generated content on their device. This selection allows them to choose the most appropriate style or genre from several options as needed, and is an important step in reflecting the user's intentions.
[0044] Content generation
[0045] The server uses an image generation algorithm to create visual content based on the analyzed text data and the style and genre selected by the user. Specific illustrations and graphics are created during this process. Similarly, a speech synthesis engine is used to generate auditory content. At this time, speech parameters are set according to the user's selection to produce natural-sounding speech and music.
[0046] Multilingual support and collaboration
[0047] Upon user request, the server can use a translation module to convert text into another language and regenerate the content. Furthermore, collaboration features are provided for sharing and co-editing the generated content with other users in real time. This allows users to efficiently advance their projects.
[0048] Output and save
[0049] The final generated visual and auditory content is displayed on the user's device, and the user can choose to download or save it. The content can be customized to the user's needs and is available across various digital media.
[0050] This system allows users to easily generate diverse digital content without requiring advanced professional skills. This significantly contributes to improving the efficiency and creativity of content creation.
[0051] The following describes the processing flow.
[0052] Step 1:
[0053] The user accesses the system via their device and logs in. Upon successful login, a text input form for content generation is displayed. The user enters the text that will be the source of the content they want to generate into the form and submits it.
[0054] Step 2:
[0055] The terminal forwards the transmitted text to the server. The server receives this text and performs tokenization, syntactic analysis, and semantic analysis of the text using a natural language processing (NLP) module. This analysis extracts the main points and context of the text.
[0056] Step 3:
[0057] Users can select the style and genre of visual and auditory content on their device. For example, options such as fantasy or scientific style, narration, and music are provided. Once selections are complete, the device sends that information to the server.
[0058] Step 4:
[0059] The server generates visual content using an image generation algorithm based on the parsed text, taking into account the selected style and genre. Sound waveforms and lighting effects are also applied in this step.
[0060] Step 5:
[0061] The server runs a speech synthesis engine to generate audio content based on text. Speech parameters such as tone, speed, and accent are adjusted according to the selected genre.
[0062] Step 6:
[0063] When a user requests multilingual content, the server uses a translation module to convert the text into the specified language. After translation, the visual and auditory content is generated again.
[0064] Step 7:
[0065] The generated content can be shared and collaboratively edited among users. The server generates a collaboration link and sends it to the user's device. Users use this link to begin collaborating with other users.
[0066] Step 8:
[0067] The final generated content is displayed on the device. Users can choose to save or download it. Following the user's instructions, the server stores the content in a database for later reuse.
[0068] (Example 1)
[0069] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0070] Currently, many digital content generation systems generate visual or auditory information based on text entered by users in natural language. However, these systems have limited multilingual support and real-time collaborative editing capabilities, failing to adequately meet the diverse needs of users. Furthermore, the inability to smoothly share and collaborate on content among different users presents inconveniences.
[0071] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0072] In this invention, the server includes means for analyzing the meaning of an input document using natural language processing technology, means for generating visual information in a selected format, means for generating auditory information in a selected category, means for translating into different languages and generating content again, and means for synchronizing the generated content with other devices and collaborating. This enables the creation of more multilingual and collaborative digital content.
[0073] "Natural language processing technology" refers to a set of methods and algorithms that enable computers to understand and analyze human language.
[0074] A "document" refers to text data entered by a user, which is the subject of natural language processing.
[0075] "Format" refers to the specific style or design pattern that a user chooses when generating visual information.
[0076] "Visual information" refers to content generated to create visual representations such as images and illustrations.
[0077] A "category" refers to a specific genre or theme that a user selects when generating auditory information.
[0078] "Auditory information" refers to content generated to create auditory expressions such as speech and music.
[0079] Translation is the process of converting a document from one language to another.
[0080] "Device" refers to hardware or software used in conjunction with a system to display and edit digital content.
[0081] "Synchronization" is the process of coordinating data across multiple devices to ensure consistency and updating information in real time.
[0082] "Collaborative work" refers to activities in which multiple users work together to edit or create a single piece of content, either simultaneously or at different times.
[0083] The system in this invention uses natural language processing technology to analyze documents input by a user and generate visual and auditory information. Specifically, a server, a terminal, and a user cooperate to carry out the digital content generation process.
[0084] The user begins the process by accessing the system using their device and logging in. After the login screen, the user is presented with a text input form. The user enters a prompt into this form, such as "Please create a story with the theme of the arrival of spring." The device then sends this input document to the server.
[0085] The server analyzes received documents using natural language processing (NLP) techniques. It performs text tokenization, syntactic analysis, and semantic analysis, and based on the results, generates visual and auditory information according to the format and category selected by the user. Visual information is created using image generation software, and auditory information is created using a speech synthesis engine.
[0086] Furthermore, the server uses translation modules as needed to convert content into other languages. This enables the generation of multilingual content. The generated content is sent from the server to the terminal in real time, and users can view the results on their terminal and download or share them.
[0087] Furthermore, the system includes collaborative features, allowing users to share generated content with other users in real time and collaboratively edit it. In this way, the invention transforms users' ideas into visual and auditory digital content, enabling efficient and flexible content creation.
[0088] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0089] Step 1:
[0090] The user accesses the system through their terminal and logs in. After logging in, a text input form is displayed. The user enters a prompt message into this form—for example, "Please create a story with the theme of the arrival of spring." The entered text is treated as the system's initial data.
[0091] Step 2:
[0092] The terminal sends the prompt message entered by the user to the server. At this time, the prompt message is encoded as a digital signal and transferred to the server according to the communication protocol. The server receives this signal, decodes it as text data, and uses it as the start data for processing.
[0093] Step 3:
[0094] The server analyzes the received text data using a natural language processing (NLP) module. Specifically, it tokenizes the input sentence and analyzes its grammatical structure using a syntactic analysis engine. Furthermore, it understands the context through semantic analysis and extracts the information necessary for generating content. The analysis results become the foundational data for subsequent data generation processes.
[0095] Step 4:
[0096] Based on the analysis results, the server generates visual and auditory information according to the format and category specified by the user. Visual information is created and output using an image generation algorithm. Auditory information is output as audio or music according to the specified audio parameters using a speech synthesis engine. Image files and audio files are generated during this process.
[0097] Step 5:
[0098] The server sends the generated visual and auditory information to the user's device. This output content is displayed on the device in real time. The user can review the generated information on their device and save or download it as needed.
[0099] Step 6:
[0100] If necessary, the user requests content generation in another language from their device. The server uses a translation module to convert existing text data into the other language. Then, it generates visual and auditory information again and sends it to the user's device. This completes the delivery of multilingual content.
[0101] Step 7:
[0102] Users can share generated information with other users within the system and collaboratively edit it. Content can be modified or added in real time via the terminal, the server reflects these changes, and the updated data is immediately sent to all participants. This enables efficient collaboration.
[0103] (Application Example 1)
[0104] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0105] In content distribution services, there is a need to provide an environment where users can easily generate and share diverse visual and auditory information, and furthermore, to realize a function that allows for real-time collaborative editing of works with other users. Existing systems have problems that hinder user convenience, such as requiring specialized technical skills for content generation and editing, and having insufficient multilingual support.
[0106] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0107] In this invention, the server includes means for analyzing input text using natural language processing technology, means for generating visual information in a specified format based on the analysis results, means for generating auditory information of a specified type based on the analysis results, and means for displaying the visual and auditory information in a virtual environment. This enables users to generate and share diverse visual and auditory information and to perform real-time collaborative editing that is compatible with other languages, even without specialized technical skills.
[0108] "Natural language processing technology" refers to the technology that enables computers to understand and process human language appropriately.
[0109] "Means of analysis" refers to a device or program that has the function of analyzing the characteristics and context of input text.
[0110] "Style" refers to the genre or style of visual information or works of art, and is associated with specific designs and methods of expression.
[0111] "Visual information" refers to images, illustrations, videos, etc., generated based on the analysis results.
[0112] "Type" refers to the category or genre of auditory information or musical works, and is associated with a specific musical style or sound profile.
[0113] "Auditory information" refers to speech, sounds, music, etc., generated based on the analysis results.
[0114] A "virtual environment" refers to a space for displaying and manipulating information on a virtual space or digital platform.
[0115] "Sharing" refers to a state in which multiple users can simultaneously view and access information or content.
[0116] "Collaborative editing" refers to a state in which multiple users can simultaneously edit and modify a single work or piece of information.
[0117] This invention provides a system that generates visual and auditory information based on text data entered by a user, and specific embodiments for implementing this system are shown below.
[0118] Users access this system using devices such as smartphones and personal computers. First, users enter text content into input forms displayed on their devices. The server then analyzes the text data using appropriate natural language processing technologies, such as the spaCy library. This analysis process includes tokenization, syntactic analysis, and semantic analysis to deepen the understanding of the language.
[0119] Based on the analyzed data, the server generates visual information in a specified format using image generation algorithms such as Stable Diffusion. It also generates auditory information of a specified type using speech synthesis engines such as Google® Text-to-Speech. The generated information is displayed in a virtual environment on a digital platform, which users can view and interact with.
[0120] Furthermore, anticipating that content will be used in multiple languages, a translation module for multilingual support (e.g., Google Translate API) has been implemented. This ensures that users can access content in different languages.
[0121] As a concrete example, if a user enters the beginning of "The Little Prince" and selects the "Fantasy" style, the server will analyze the entered text and generate a fantastical illustration and audio narration that matches the story. An example of a prompt in this case might be, "Please create fantasy-style illustrations and audio content from the text."
[0122] This system allows users to create and share engaging digital content without requiring specialized technical skills.
[0123] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0124] Step 1:
[0125] The user accesses the system using a terminal and logs in. The user enters the text they want to generate into a text input form on the terminal. This input data is sent to the server as basic information for subsequent analysis and generation.
[0126] Step 2:
[0127] The server receives text data sent by the user. Based on the received data, it starts analysis using natural language processing (NLP) techniques. Specifically, it tokenizes the text using an NLP library such as spaCy, and then performs syntactic and semantic analysis. This process captures the characteristics and context of the text, and outputs the basic data for the next generation step.
[0128] Step 3:
[0129] Users select the style and genre of content to be generated on their device. This selection is sent to the server and reflected in the generation process. This selection influences the algorithms used to generate visual and auditory content.
[0130] Step 4:
[0131] The server generates visual information using image generation algorithms such as Stable Diffusion, based on the analyzed text data and user selection information. In this step, the analyzed data and format are used as input, and the output is a graphic in the specified style. The generated visual information is temporarily stored on the server.
[0132] Step 5:
[0133] In parallel, the server uses a text-to-speech engine, such as Google Text-to-Speech, to generate auditory information based on the parsed text data. The parsed data and type are used as input, and the output is audio content of the specified type. This generates natural-sounding speech that reflects the user's intent.
[0134] Step 6:
[0135] The server generates content in multiple languages, using translation modules (e.g., Google Translate API) as needed. In this step, the original text and the user-selected language are used as input, and the output is the translated text and the corresponding audio content.
[0136] Step 7:
[0137] The generated visual and auditory information is sent to the user's device and displayed in the virtual environment. The user can view and interact with this information. Furthermore, the displayed content can be shared and collaboratively edited with other users, facilitating real-time communication and creative work.
[0138] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0139] The present invention is a visual and auditory digital content generation system that combines an emotion engine that recognizes user emotions. Embodiments of this system will be described below.
[0140] Overall structure and user interface
[0141] Users log in to the system via their device and enter the necessary information into a text input form. This entered text becomes the basic data for the generated digital content. Users can then select visual styles and auditory genres to enjoy a seamless experience.
[0142] Emotion recognition and data analysis
[0143] The server analyzes the text sent from the terminal using natural language processing techniques. During this analysis, an emotion engine recognizes the user's emotions and determines the emotional tone based on the text. The recognized emotions are then used to dynamically adjust the subsequent content generation process.
[0144] Content generation process
[0145] Based on the analysis results and the emotion engine's output, the server generates visual digital content that takes style selection and emotional tone into consideration. For example, if the user is excited, a colorful and dynamic design will be chosen. Similarly, auditory content is generated based on emotion, with the speech synthesis engine adjusting the tone and tempo.
[0146] Multilingual support and collaboration features
[0147] The server's multilingual translation feature provides users with the option to generate content in different languages. The generated content can be shared with other users in real time, and collaboration features facilitate co-editing. This feature is particularly useful when project teams are multinational.
[0148] Content output and saving
[0149] The generated visual and auditory content is displayed on the user's device. The user can review the content and choose to download it or save it on the server as needed. The server stores the content in a database accordingly and manages it as a history.
[0150] In this way, systems that combine emotion engines provide personalized digital content that reflects the user's emotional state, aiming to improve creativity and efficiency. For example, if a user inputs a part of an emotionally moving story, digital content is generated consisting of a visual design with calming colors and emotionally moving music.
[0151] The following describes the processing flow.
[0152] Step 1:
[0153] The user logs into the system using their device and accesses the dashboard. They enter the text that will be the source of the content they want to generate into the displayed text input form and click the "Submit" button.
[0154] Step 2:
[0155] The terminal sends the text data entered by the user to the emotion engine. The server receives this text and analyzes its content using natural language processing techniques.
[0156] Step 3:
[0157] The server extracts contextual information from the analyzed text and identifies the user's emotions through an emotion engine. For example, the emotion engine determines whether the text contains emotions such as joy, sadness, or excitement.
[0158] Step 4:
[0159] The user receives an option on their device to select their preferred visual style and auditory genre. For example, they might choose "animation style" or "classical music."
[0160] Step 5:
[0161] The server begins generating visual content based on the emotions recognized by the emotion engine and the user's selections. The server uses an image generation algorithm to dynamically determine colors and designs that correspond to the emotions.
[0162] Step 6:
[0163] Similarly, the server operates a speech synthesis engine to generate auditory content. Here, the tone and tempo of the voice are adjusted to match the user's emotions, creating a sound that fits those emotions.
[0164] Step 7:
[0165] If the user has selected multilingual support, the server will use a translation API to convert the text into other languages and regenerate the content as needed.
[0166] Step 8:
[0167] The generated visual and auditory content is displayed on the device. Users can review the results and, if satisfied, download the content or share it with others using collaboration tools.
[0168] Step 9:
[0169] If a user chooses to save content, the server stores it in a database and manages it as part of the user's project history. This allows for later reuse and editing.
[0170] (Example 2)
[0171] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0172] A challenge in modern digital content creation is the difficulty in generating content that fully reflects the user's emotions and intentions. Conventional generation methods are limited to simple conversions based on input text data, making it difficult to provide customized content tailored to the specific emotions or situations expressed by the user. Furthermore, support for content sharing and collaborative editing across different languages is insufficient.
[0173] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0174] In this invention, the server includes means for performing sentiment analysis to recognize the user's emotions, means for dynamically adjusting and generating visual and auditory digital content based on the emotions, and means for generating customized content that reflects the user's input information and emotions. This enables the provision of personalized digital content that enhances the user experience.
[0175] A "user" is a person or group that operates the system and provides input data.
[0176] "Sentiment analysis" is the process of detecting a user's emotional state from input text data and interpreting the data based on that.
[0177] "Visual digital content" refers to content that includes images, videos, or visual elements generated in a digital format.
[0178] "Auditory digital content" refers to content that generates sound-related elements, such as voices and music, in a digital format.
[0179] "Natural language processing" is the technology that enables computers to understand, interpret, and manipulate human language.
[0180] "Multilingual translation" is the process of translating data between different languages and providing the translated result in a specific language.
[0181] "Collaborative work tools" are features that allow multiple users to jointly edit or share content.
[0182] A "database" is a system for organizing, efficiently storing, searching, and managing data.
[0183] This system is designed to allow users to generate digital content via their devices. Users log in using an interface on their devices and then input text data. This text data is sent to the server and used as the basis for the digital content.
[0184] The server uses NLP libraries and frameworks (e.g., SpaCy, NLTK) to analyze this text data using natural language processing techniques. Based on this analysis, it runs an emotion analysis engine to recognize the user's emotions. The server then uses generative AI models (e.g., DALL-E, VQ-VAE-2) to generate visual and auditory digital content. During this process, a connected speech synthesis engine adjusts the tone and tempo of the generated audio data.
[0185] The server uses a multilingual translation API to translate generated content into different languages, enabling the provision of content that supports various language environments. Furthermore, the generated content is designed to be shared and collaboratively edited with other users in real time. For this purpose, it incorporates collaborative work tools, allowing for smooth operation even when project teams are composed of members from multiple countries.
[0186] The content is displayed on the user's device, and after viewing it, the user can download or save it to the server as needed. The server manages the content in a database and stores it as a history so that the user can access it later.
[0187] For example, if a user enters the prompt "Generate a heartwarming story," the server will generate digital content including a visual design with gentle, calming colors and moving music. In this way, it becomes possible to provide personalized digital content based on the user's intentions and emotions.
[0188] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0189] Step 1:
[0190] Users log in to the system via their device. During login, the user enters authentication information, which the device then sends to the server. The server verifies this information to confirm that the user is a legitimate user. If the login is successful, the user can access the system.
[0191] Step 2:
[0192] The user enters text data on the terminal interface. In the input form, the user provides information related to the digital content in text format. The terminal collects this input data and sends it to the server. This entered text data is used as the basis for the digital content to be generated.
[0193] Step 3:
[0194] The server performs natural language processing on the received text data. The specific software used is an NLP library such as SpaCy or NLTK, which performs grammatical and semantic analysis. The input data is analyzed, and as a result, the text's structure and semantic information are extracted. This information obtained from the analysis is then used for subsequent processing.
[0195] Step 4:
[0196] The server uses the analysis results to perform sentiment analysis. This process extracts the emotions contained in the input text using a sentiment analysis engine. The analysis results are then used during subsequent content generation, forming the basis for data adjustments based on the user's emotional state.
[0197] Step 5:
[0198] The server uses generative AI models to generate visual and auditory digital content. Specifically, it uses generative AI models such as DALL-E and VQ-VAE-2 to create content that matches the user's emotions. In this process, the design's color scheme and audio tone are set based on the emotional tone. The output content is then provided to the user.
[0199] Step 6:
[0200] The server uses a multilingual translation API to translate the generated content into other languages. This makes the content understandable to users in multilingual environments. The server processes the translated data in real time and generates output that includes corresponding language information.
[0201] Step 7:
[0202] The server sends the generated content to the user's device and simultaneously generates a link for sharing with other users. Users can use this link in real time to collaborate on editing with other users. The generated link information is passed around via email, chat, etc.
[0203] Step 8:
[0204] The user reviews the content received on their device and downloads or saves it to the server as needed. The server, following the user's instructions, stores the content in a database and manages it as a history. This makes it easy for the user to access the content at a later date.
[0205] (Application Example 2)
[0206] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0207] In the modern era, digital content is delivered in a variety of formats, but there is a challenge in the lack of dynamic content adjustments that respond to user emotions. Furthermore, there is a demand for content that is more personalized and relatable to viewers. To achieve this, technology is needed that accurately analyzes user emotions and dynamically generates and adjusts content based on those results.
[0208] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0209] In this invention, the server includes means for analyzing the meaning of input text using natural language processing technology, means for analyzing the user's emotions using emotion recognition technology and dynamically adjusting visual and auditory digital content according to those emotions, and means for changing the tone and theme of a video according to the user's emotions. This makes it possible to provide personalized digital content based on the user's emotions.
[0210] "Natural language processing technology" is a technology that uses computers to understand and analyze human language.
[0211] "Text semantic analysis" involves structurally analyzing input text to understand its content and intent.
[0212] "Generating visual digital content" means creating media that can be displayed in a digital format based on visual information.
[0213] "Generating auditory digital content" means recording or generating auditory information in a digital format to create playable audio data.
[0214] "Emotion recognition technology" is a technology that detects and analyzes the emotions expressed by a user as data.
[0215] "Dynamically adjusting content" means changing the nature or structure of content in real time or as needed.
[0216] "Changing the tone or theme of a video" means altering the atmosphere or message of the video to provide viewers with a different visual experience.
[0217] This invention is a system that allows users to access the system via a terminal to generate personalized visual and auditory digital content based on their emotions. Specific embodiments thereof are described below.
[0218] The server first analyzes the text data received from the user using natural language processing techniques. This analysis extracts the user's intentions and emotions from the text. Specifically, Python's natural language processing library (e.g., NLTK) may be used for this analysis.
[0219] Based on the analysis results, the server then uses emotion recognition technology to understand the user's emotions in detail. In this process, emotion recognition technology identifies the user's emotions and determines what tone and theme of content is appropriate.
[0220] Based on the emotional information received, the server generates visual and auditory digital content. The visual content is dynamically styled, with changes in color tone and brightness as needed. The auditory content, on the other hand, uses a speech synthesis engine to apply emotionally appropriate tone and tempo.
[0221] For example, a prompt message given to the system might be, "Design the best visual content and music for when the user wants to relax." Based on this instruction, the system generates and provides a video of a calm nighttime landscape and soothing music.
[0222] Finally, the generated content is sent to the user's device and displayed or played. At this point, the user can review the generated content and determine if it is appropriate. Furthermore, the content can be generated in multiple languages and can be shared and collaboratively edited among users. The server records the generation process and stores it as data for future content generation.
[0223] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0224] Step 1:
[0225] The terminal receives text data from the user as input. The user uses prompts to specify the intent and content of the material they want to generate. This input text data is sent to the server as initial system data.
[0226] Step 2:
[0227] The server analyzes the received text data using natural language processing techniques. Specifically, it extracts nouns, verbs, and other elements from the input text using a natural language processing library to understand the context and sentiment. The output of this analysis contains information about the user's intent and sentiment, and serves as foundational data for the next step.
[0228] Step 3:
[0229] The server applies emotion recognition technology based on the analysis results to identify the user's emotions. Based on the analysis data received as input, the emotion engine determines the intensity and type of emotion. For example, emotions such as joy and sadness are extracted, and the results influence subsequent content generation.
[0230] Step 4:
[0231] The server takes identified emotional information as input and generates visual and auditory digital content that matches the emotion. Visual content is dynamically adjusted in terms of color and layout, while auditory content has its tone and tempo set using a speech synthesis engine. The output is customized digital content that matches the user's emotion.
[0232] Step 5:
[0233] The server sends the generated content to the terminal. The user can experience the output visually or aurally and judge the appropriateness of the content.
[0234] Step 6:
[0235] The server records the generated content and its generation process, and stores it in a database. This information is used to improve the content generation process in the future and to understand user preferences.
[0236] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0237] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0238] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0239] [Second Embodiment]
[0240] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0241] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0242] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0243] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0244] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0245] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0246] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0247] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0248] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0249] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0250] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0251] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0252] This invention provides users with a system for generating visual and auditory digital content using natural language processing technology. The following is an example of this system.
[0253] User interface and text input
[0254] The user accesses the system using their device and logs in to begin the content generation process. A text input form appears on the device, and the user enters text related to the content they want to generate into this form. For example, this could be a portion of a novel or the beginning of a blog post.
[0255] Data reception and analysis
[0256] The server receives text data sent from the user's terminal. The received text is parsed by the server's natural language processing (NLP) module. Specifically, it performs tokenization, syntactic analysis, and semantic analysis of the text to determine what visual and auditory content is appropriate based on its characteristics and context.
[0257] Style and Genre Selection
[0258] Users can also select the style and genre of the generated content on their device. This selection allows them to choose the most appropriate style or genre from several options as needed, and is an important step in reflecting the user's intentions.
[0259] Content generation
[0260] The server uses an image generation algorithm to create visual content based on the analyzed text data and the style and genre selected by the user. Specific illustrations and graphics are created during this process. Similarly, a speech synthesis engine is used to generate auditory content. At this time, speech parameters are set according to the user's selection to produce natural-sounding speech and music.
[0261] Multilingual support and collaboration
[0262] Upon user request, the server can use a translation module to convert text into another language and regenerate the content. Furthermore, collaboration features are provided for sharing and co-editing the generated content with other users in real time. This allows users to efficiently advance their projects.
[0263] Output and save
[0264] The final generated visual and auditory content is displayed on the user's device, and the user can choose to download or save it. The content can be customized to the user's needs and is available across various digital media.
[0265] This system allows users to easily generate diverse digital content without requiring advanced professional skills. This significantly contributes to improving the efficiency and creativity of content creation.
[0266] The following describes the processing flow.
[0267] Step 1:
[0268] The user accesses the system via their device and logs in. Upon successful login, a text input form for content generation is displayed. The user enters the text that will be the source of the content they want to generate into the form and submits it.
[0269] Step 2:
[0270] The terminal forwards the transmitted text to the server. The server receives this text and performs tokenization, syntactic analysis, and semantic analysis of the text using a natural language processing (NLP) module. This analysis extracts the main points and context of the text.
[0271] Step 3:
[0272] Users can select the style and genre of visual and auditory content on their device. For example, options such as fantasy or scientific style, narration, and music are provided. Once selections are complete, the device sends that information to the server.
[0273] Step 4:
[0274] The server generates visual content using an image generation algorithm based on the analyzed text, taking into account the selected style and genre. Waveforms of the home and lighting effects are also applied in this step.
[0275] Step 5:
[0276] The server operates a speech synthesis engine to generate auditory content based on the text. The speech parameters such as tone, speed, and accent of the voice are adjusted according to the selected genre.
[0277] Step 6:
[0278] If the user requests multilingual content, the server uses a translation module to convert the text into the specified language. After translation, visual and auditory content generation is performed again.
[0279] Step 7:
[0280] The generated content can be shared and co - edited among users. The server generates a link for collaboration and sends it to the terminal. The user uses this link to start a collaborative work with other users.
[0281] Step 8:
[0282] Finally, the generated content is displayed on the terminal. The user can select options for saving or downloading. According to the user's instructions, the server stores the content in the database so that it can be reused later.
[0283] (Example 1)
[0284] Next, Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0285] Currently, many digital content generation systems generate visual or auditory information based on the text input by users in natural language. However, in these systems, multilingual support and real-time collaborative editing functions are limited, and they cannot fully meet the diverse needs of users. In addition, there are also inconveniences caused by the fact that content sharing and collaboration between different users cannot be smoothly carried out.
[0286] The specific processing by the specific processing unit 290 of the data processing device 12 in Embodiment 1 is realized by the following means.
[0287] In this invention, the server includes means for analyzing the meaning of a document input using natural language processing technology, means for generating visual information in a selected format, means for generating auditory information in a selected category, means for performing translation in different languages and generating the content again, and means for synchronizing the generated content with other devices and performing collaborative work. Thereby, it becomes possible to generate more multilingual and highly collaborative digital content.
[0288] "Natural language processing technology" refers to a series of methods and algorithms for a computer to understand and analyze human language.
[0289] "Document" refers to the text data input by the user and is the object of natural language processing.
[0290] "Format" refers to a specific style or design pattern selected by the user in visual information generation.
[0291] "Visual information" refers to the content generated to create visual expressions such as images and illustrations.
[0292] "Category" refers to a specific genre or theme selected by the user in auditory information generation.
[0293] "Auditory information" refers to content generated to create auditory expressions such as speech and music.
[0294] Translation is the process of converting a document from one language to another.
[0295] "Device" refers to hardware or software used in conjunction with a system to display and edit digital content.
[0296] "Synchronization" is the process of coordinating data across multiple devices to ensure consistency and updating information in real time.
[0297] "Collaborative work" refers to activities in which multiple users work together to edit or create a single piece of content, either simultaneously or at different times.
[0298] The system in this invention uses natural language processing technology to analyze documents input by a user and generate visual and auditory information. Specifically, a server, a terminal, and a user cooperate to carry out the digital content generation process.
[0299] The user begins the process by accessing the system using their device and logging in. After the login screen, the user is presented with a text input form. The user enters a prompt into this form, such as "Please create a story with the theme of the arrival of spring." The device then sends this input document to the server.
[0300] The server analyzes received documents using natural language processing (NLP) techniques. It performs text tokenization, syntactic analysis, and semantic analysis, and based on the results, generates visual and auditory information according to the format and category selected by the user. Visual information is created using image generation software, and auditory information is created using a speech synthesis engine.
[0301] In addition, the server uses a translation module as needed to perform conversion into other languages. This enables the generation of content in multiple languages. The generated content is transmitted in real time from the server to the terminal, and the user can check the results on the terminal and perform downloads and sharing.
[0302] Furthermore, the system has a collaborative work function, and users can share the generated content with other users in real time and cooperate in editing. In this way, this invention gives form to the ideas of users as visual and auditory digital content, realizing efficient and flexible content generation.
[0303] The flow of specific processing in Example 1 will be described using FIG. 11.
[0304] Step 1:
[0305] The user accesses the system through the terminal and logs in. After logging in, a text input form is displayed. The user inputs a prompt sentence, such as "Please create a story themed on the arrival of spring," into this form. The input text is treated as the initial data of the system.
[0306] Step 2:
[0307] The terminal transmits the prompt sentence input by the user to the server. At this time, the prompt sentence is encoded as a digital signal and transferred to the server according to the communication protocol. The server receives this signal, decodes it as text data, and uses it as the start data for processing.
[0308] Step 3:
[0309] The server analyzes the received text data using a natural language processing (NLP) module. Specifically, it tokenizes the input sentence and analyzes its grammatical structure using a syntactic analysis engine. Furthermore, it understands the context through semantic analysis and extracts the information necessary for generating content. The analysis results become the foundational data for subsequent data generation processes.
[0310] Step 4:
[0311] Based on the analysis results, the server generates visual and auditory information according to the format and category specified by the user. Visual information is created and output using an image generation algorithm. Auditory information is output as audio or music according to the specified audio parameters using a speech synthesis engine. Image files and audio files are generated during this process.
[0312] Step 5:
[0313] The server sends the generated visual and auditory information to the user's device. This output content is displayed on the device in real time. The user can review the generated information on their device and save or download it as needed.
[0314] Step 6:
[0315] If necessary, the user requests content generation in another language from their device. The server uses a translation module to convert existing text data into the other language. Then, it generates visual and auditory information again and sends it to the user's device. This completes the delivery of multilingual content.
[0316] Step 7:
[0317] Users can share generated information with other users within the system and collaboratively edit it. Content can be modified or added in real time via the terminal, the server reflects these changes, and the updated data is immediately sent to all participants. This enables efficient collaboration.
[0318] (Application Example 1)
[0319] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0320] In content distribution services, there is a need to provide an environment where users can easily generate and share diverse visual and auditory information, and furthermore, to realize a function that allows for real-time collaborative editing of works with other users. Existing systems have problems that hinder user convenience, such as requiring specialized technical skills for content generation and editing, and having insufficient multilingual support.
[0321] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0322] In this invention, the server includes means for analyzing input text using natural language processing technology, means for generating visual information in a specified format based on the analysis results, means for generating auditory information of a specified type based on the analysis results, and means for displaying the visual and auditory information in a virtual environment. This enables users to generate and share diverse visual and auditory information and to perform real-time collaborative editing that is compatible with other languages, even without specialized technical skills.
[0323] "Natural language processing technology" refers to the technology that enables computers to understand and process human language appropriately.
[0324] "Means of analysis" refers to a device or program that has the function of analyzing the characteristics and context of input text.
[0325] "Style" refers to the genre or style of visual information or works of art, and is associated with specific designs and methods of expression.
[0326] "Visual information" refers to images, illustrations, videos, etc., generated based on the analysis results.
[0327] "Type" refers to the category or genre of auditory information or musical works, and is associated with a specific musical style or sound profile.
[0328] "Auditory information" refers to speech, sounds, music, etc., generated based on the analysis results.
[0329] A "virtual environment" refers to a space for displaying and manipulating information on a virtual space or digital platform.
[0330] "Sharing" refers to a state in which multiple users can simultaneously view and access information or content.
[0331] "Collaborative editing" refers to a state in which multiple users can simultaneously edit and modify a single work or piece of information.
[0332] This invention provides a system that generates visual and auditory information based on text data entered by a user, and specific embodiments for implementing this system are shown below.
[0333] Users access this system using devices such as smartphones and personal computers. First, users enter text content into input forms displayed on their devices. The server then analyzes the text data using appropriate natural language processing technologies, such as the spaCy library. This analysis process includes tokenization, syntactic analysis, and semantic analysis to deepen the understanding of the language.
[0334] Based on the analyzed data, the server generates visual information in a specified format using image generation algorithms such as Stable Diffusion. It also generates auditory information of a specified type using a speech synthesis engine such as Google Text-to-Speech. The generated information is displayed in a virtual environment on a digital platform, which users can view and interact with.
[0335] Furthermore, anticipating that content will be used in multiple languages, a translation module for multilingual support (e.g., Google Translate API) has been implemented. This ensures that users can access content in different languages.
[0336] As a concrete example, if a user enters the beginning of "The Little Prince" and selects the "Fantasy" style, the server will analyze the entered text and generate a fantastical illustration and audio narration that matches the story. An example of a prompt in this case might be, "Please create fantasy-style illustrations and audio content from the text."
[0337] This system allows users to create and share engaging digital content without requiring specialized technical skills.
[0338] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0339] Step 1:
[0340] The user accesses the system using a terminal and logs in. The user enters the text they want to generate into a text input form on the terminal. This input data is sent to the server as basic information for subsequent analysis and generation.
[0341] Step 2:
[0342] The server receives text data sent by the user. Based on the received data, it starts analysis using natural language processing (NLP) techniques. Specifically, it tokenizes the text using an NLP library such as spaCy, and then performs syntactic and semantic analysis. This process captures the characteristics and context of the text, and outputs the basic data for the next generation step.
[0343] Step 3:
[0344] Users select the style and genre of content to be generated on their device. This selection is sent to the server and reflected in the generation process. This selection influences the algorithms used to generate visual and auditory content.
[0345] Step 4:
[0346] The server generates visual information using image generation algorithms such as Stable Diffusion, based on the analyzed text data and user selection information. In this step, the analyzed data and format are used as input, and the output is a graphic in the specified style. The generated visual information is temporarily stored on the server.
[0347] Step 5:
[0348] In parallel, the server uses a text-to-speech engine, such as Google Text-to-Speech, to generate auditory information based on the parsed text data. The parsed data and type are used as input, and the output is audio content of the specified type. This generates natural-sounding speech that reflects the user's intent.
[0349] Step 6:
[0350] The server generates content in multiple languages, using translation modules (e.g., Google Translate API) as needed. In this step, the original text and the user-selected language are used as input, and the output is the translated text and the corresponding audio content.
[0351] Step 7:
[0352] The generated visual and auditory information is sent to the user's device and displayed in the virtual environment. The user can view and interact with this information. Furthermore, the displayed content can be shared and collaboratively edited with other users, facilitating real-time communication and creative work.
[0353] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0354] The present invention is a visual and auditory digital content generation system that combines an emotion engine that recognizes user emotions. Embodiments of this system will be described below.
[0355] Overall structure and user interface
[0356] Users log in to the system via their device and enter the necessary information into a text input form. This entered text becomes the basic data for the generated digital content. Users can then select visual styles and auditory genres to enjoy a seamless experience.
[0357] Emotion recognition and data analysis
[0358] The server analyzes the text sent from the terminal using natural language processing techniques. During this analysis, an emotion engine recognizes the user's emotions and determines the emotional tone based on the text. The recognized emotions are then used to dynamically adjust the subsequent content generation process.
[0359] Content generation process
[0360] Based on the analysis results and the emotion engine's output, the server generates visual digital content that takes style selection and emotional tone into consideration. For example, if the user is excited, a colorful and dynamic design will be chosen. Similarly, auditory content is generated based on emotion, with the speech synthesis engine adjusting the tone and tempo.
[0361] Multilingual support and collaboration features
[0362] The server's multilingual translation feature provides users with the option to generate content in different languages. The generated content can be shared with other users in real time, and collaboration features facilitate co-editing. This feature is particularly useful when project teams are multinational.
[0363] Content output and saving
[0364] The generated visual and auditory content is displayed on the user's device. The user can review the content and choose to download it or save it on the server as needed. The server stores the content in a database accordingly and manages it as a history.
[0365] In this way, systems that combine emotion engines provide personalized digital content that reflects the user's emotional state, aiming to improve creativity and efficiency. For example, if a user inputs a part of an emotionally moving story, digital content is generated consisting of a visual design with calming colors and emotionally moving music.
[0366] The following describes the processing flow.
[0367] Step 1:
[0368] The user logs into the system using their device and accesses the dashboard. They enter the text that will be the source of the content they want to generate into the displayed text input form and click the "Submit" button.
[0369] Step 2:
[0370] The terminal sends the text data entered by the user to the emotion engine. The server receives this text and analyzes its content using natural language processing techniques.
[0371] Step 3:
[0372] The server extracts contextual information from the analyzed text and identifies the user's emotions through an emotion engine. For example, the emotion engine determines whether the text contains emotions such as joy, sadness, or excitement.
[0373] Step 4:
[0374] The user receives an option on their device to select their preferred visual style and auditory genre. For example, they might choose "animation style" or "classical music."
[0375] Step 5:
[0376] The server begins generating visual content based on the emotions recognized by the emotion engine and the user's selections. The server uses an image generation algorithm to dynamically determine colors and designs that correspond to the emotions.
[0377] Step 6:
[0378] Similarly, the server operates a speech synthesis engine to generate auditory content. Here, the tone and tempo of the voice are adjusted to match the user's emotions, creating a sound that fits those emotions.
[0379] Step 7:
[0380] If the user has selected multilingual support, the server will use a translation API to convert the text into other languages and regenerate the content as needed.
[0381] Step 8:
[0382] The generated visual and auditory content is displayed on the device. Users can review the results and, if satisfied, download the content or share it with others using collaboration tools.
[0383] Step 9:
[0384] If a user chooses to save content, the server stores it in a database and manages it as part of the user's project history. This allows for later reuse and editing.
[0385] (Example 2)
[0386] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0387] A challenge in modern digital content creation is the difficulty in generating content that fully reflects the user's emotions and intentions. Conventional generation methods are limited to simple conversions based on input text data, making it difficult to provide customized content tailored to the specific emotions or situations expressed by the user. Furthermore, support for content sharing and collaborative editing across different languages is insufficient.
[0388] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0389] In this invention, the server includes means for performing sentiment analysis to recognize the user's emotions, means for dynamically adjusting and generating visual and auditory digital content based on the emotions, and means for generating customized content that reflects the user's input information and emotions. This enables the provision of personalized digital content that enhances the user experience.
[0390] A "user" is a person or group that operates the system and provides input data.
[0391] "Sentiment analysis" is the process of detecting a user's emotional state from input text data and interpreting the data based on that.
[0392] "Visual digital content" refers to content that includes images, videos, or visual elements generated in a digital format.
[0393] "Auditory digital content" refers to content that generates sound-related elements, such as voices and music, in a digital format.
[0394] "Natural language processing" is the technology that enables computers to understand, interpret, and manipulate human language.
[0395] "Multilingual translation" is the process of translating data between different languages and providing the translated result in a specific language.
[0396] "Collaborative work tools" are features that allow multiple users to jointly edit or share content.
[0397] A "database" is a system for organizing, efficiently storing, searching, and managing data.
[0398] This system is designed to allow users to generate digital content via their devices. Users log in using an interface on their devices and then input text data. This text data is sent to the server and used as the basis for the digital content.
[0399] The server uses NLP libraries and frameworks (e.g., SpaCy, NLTK) to analyze this text data using natural language processing techniques. Based on this analysis, it runs an emotion analysis engine to recognize the user's emotions. The server then uses generative AI models (e.g., DALL-E, VQ-VAE-2) to generate visual and auditory digital content. During this process, a connected speech synthesis engine adjusts the tone and tempo of the generated audio data.
[0400] The server uses a multilingual translation API to translate generated content into different languages, enabling the provision of content that supports various language environments. Furthermore, the generated content is designed to be shared and collaboratively edited with other users in real time. For this purpose, it incorporates collaborative work tools, allowing for smooth operation even when project teams are composed of members from multiple countries.
[0401] The content is displayed on the user's device, and after viewing it, the user can download or save it to the server as needed. The server manages the content in a database and stores it as a history so that the user can access it later.
[0402] For example, if a user enters the prompt "Generate a heartwarming story," the server will generate digital content including a visual design with gentle, calming colors and moving music. In this way, it becomes possible to provide personalized digital content based on the user's intentions and emotions.
[0403] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0404] Step 1:
[0405] Users log in to the system via their device. During login, the user enters authentication information, which the device then sends to the server. The server verifies this information to confirm that the user is a legitimate user. If the login is successful, the user can access the system.
[0406] Step 2:
[0407] The user enters text data on the terminal interface. In the input form, the user provides information related to the digital content in text format. The terminal collects this input data and sends it to the server. This entered text data is used as the basis for the digital content to be generated.
[0408] Step 3:
[0409] The server performs natural language processing on the received text data. The specific software used is an NLP library such as SpaCy or NLTK, which performs grammatical and semantic analysis. The input data is analyzed, and as a result, the text's structure and semantic information are extracted. This information obtained from the analysis is then used for subsequent processing.
[0410] Step 4:
[0411] The server uses the analysis results to perform sentiment analysis. This process extracts the emotions contained in the input text using a sentiment analysis engine. The analysis results are then used during subsequent content generation, forming the basis for data adjustments based on the user's emotional state.
[0412] Step 5:
[0413] The server uses generative AI models to generate visual and auditory digital content. Specifically, it uses generative AI models such as DALL-E and VQ-VAE-2 to create content that matches the user's emotions. In this process, the design's color scheme and audio tone are set based on the emotional tone. The output content is then provided to the user.
[0414] Step 6:
[0415] The server uses a multilingual translation API to translate the generated content into other languages. This makes the content understandable to users in multilingual environments. The server processes the translated data in real time and generates output that includes corresponding language information.
[0416] Step 7:
[0417] The server sends the generated content to the user's device and simultaneously generates a link for sharing with other users. Users can use this link in real time to collaborate on editing with other users. The generated link information is passed around via email, chat, etc.
[0418] Step 8:
[0419] The user reviews the content received on their device and downloads or saves it to the server as needed. The server, following the user's instructions, stores the content in a database and manages it as a history. This makes it easy for the user to access the content at a later date.
[0420] (Application Example 2)
[0421] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0422] In the modern era, digital content is delivered in a variety of formats, but there is a challenge in the lack of dynamic content adjustments that respond to user emotions. Furthermore, there is a demand for content that is more personalized and relatable to viewers. To achieve this, technology is needed that accurately analyzes user emotions and dynamically generates and adjusts content based on those results.
[0423] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0424] In this invention, the server includes means for analyzing the meaning of input text using natural language processing technology, means for analyzing the user's emotions using emotion recognition technology and dynamically adjusting visual and auditory digital content according to those emotions, and means for changing the tone and theme of a video according to the user's emotions. This makes it possible to provide personalized digital content based on the user's emotions.
[0425] "Natural language processing technology" is a technology that uses computers to understand and analyze human language.
[0426] "Text semantic analysis" involves structurally analyzing input text to understand its content and intent.
[0427] "Generating visual digital content" means creating media that can be displayed in a digital format based on visual information.
[0428] "Generating auditory digital content" means recording or generating auditory information in a digital format to create playable audio data.
[0429] "Emotion recognition technology" is a technology that detects and analyzes the emotions expressed by a user as data.
[0430] "Dynamically adjusting content" means changing the nature or structure of content in real time or as needed.
[0431] "Changing the tone or theme of a video" means altering the atmosphere or message of the video to provide viewers with a different visual experience.
[0432] This invention is a system that allows users to access the system via a terminal to generate personalized visual and auditory digital content based on their emotions. Specific embodiments thereof are described below.
[0433] The server first analyzes the text data received from the user using natural language processing techniques. This analysis extracts the user's intentions and emotions from the text. Specifically, Python's natural language processing library (e.g., NLTK) may be used for this analysis.
[0434] Based on the analysis results, the server then uses emotion recognition technology to understand the user's emotions in detail. In this process, emotion recognition technology identifies the user's emotions and determines what tone and theme of content is appropriate.
[0435] Based on the emotional information received, the server generates visual and auditory digital content. The visual content is dynamically styled, with changes in color tone and brightness as needed. The auditory content, on the other hand, uses a speech synthesis engine to apply emotionally appropriate tone and tempo.
[0436] For example, a prompt message given to the system might be, "Design the best visual content and music for when the user wants to relax." Based on this instruction, the system generates and provides a video of a calm nighttime landscape and soothing music.
[0437] Finally, the generated content is sent to the user's device and displayed or played. At this point, the user can review the generated content and determine if it is appropriate. Furthermore, the content can be generated in multiple languages and can be shared and collaboratively edited among users. The server records the generation process and stores it as data for future content generation.
[0438] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0439] Step 1:
[0440] The terminal receives text data from the user as input. The user uses prompts to specify the intent and content of the material they want to generate. This input text data is sent to the server as initial system data.
[0441] Step 2:
[0442] The server analyzes the received text data using natural language processing techniques. Specifically, it extracts nouns, verbs, and other elements from the input text using a natural language processing library to understand the context and sentiment. The output of this analysis contains information about the user's intent and sentiment, and serves as foundational data for the next step.
[0443] Step 3:
[0444] The server applies emotion recognition technology based on the analysis results to identify the user's emotions. Based on the analysis data received as input, the emotion engine determines the intensity and type of emotion. For example, emotions such as joy and sadness are extracted, and the results influence subsequent content generation.
[0445] Step 4:
[0446] The server takes identified emotional information as input and generates visual and auditory digital content that matches the emotion. Visual content is dynamically adjusted in terms of color and layout, while auditory content has its tone and tempo set using a speech synthesis engine. The output is customized digital content that matches the user's emotion.
[0447] Step 5:
[0448] The server sends the generated content to the terminal. The user can experience the output visually or aurally and judge the appropriateness of the content.
[0449] Step 6:
[0450] The server records the generated content and its generation process, and stores it in a database. This information is used to improve the content generation process in the future and to understand user preferences.
[0451] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0452] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0453] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0454] [Third Embodiment]
[0455] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0456] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0457] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0458] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0459] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0460] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0461] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0462] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0463] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0464] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0465] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0466] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0467] This invention provides users with a system for generating visual and auditory digital content using natural language processing technology. The following is an example of this system.
[0468] User interface and text input
[0469] The user accesses the system using their device and logs in to begin the content generation process. A text input form appears on the device, and the user enters text related to the content they want to generate into this form. For example, this could be a portion of a novel or the beginning of a blog post.
[0470] Data reception and analysis
[0471] The server receives text data sent from the user's terminal. The received text is parsed by the server's natural language processing (NLP) module. Specifically, it performs tokenization, syntactic analysis, and semantic analysis of the text to determine what visual and auditory content is appropriate based on its characteristics and context.
[0472] Style and Genre Selection
[0473] Users can also select the style and genre of the generated content on their device. This selection allows them to choose the most appropriate style or genre from several options as needed, and is an important step in reflecting the user's intentions.
[0474] Content generation
[0475] The server uses an image generation algorithm to create visual content based on the analyzed text data and the style and genre selected by the user. Specific illustrations and graphics are created during this process. Similarly, a speech synthesis engine is used to generate auditory content. At this time, speech parameters are set according to the user's selection to produce natural-sounding speech and music.
[0476] Multilingual support and collaboration
[0477] Upon user request, the server can use a translation module to convert text into another language and regenerate the content. Furthermore, collaboration features are provided for sharing and co-editing the generated content with other users in real time. This allows users to efficiently advance their projects.
[0478] Output and save
[0479] The final generated visual and auditory content is displayed on the user's device, and the user can choose to download or save it. The content can be customized to the user's needs and is available across various digital media.
[0480] This system allows users to easily generate diverse digital content without requiring advanced professional skills. This significantly contributes to improving the efficiency and creativity of content creation.
[0481] The following describes the processing flow.
[0482] Step 1:
[0483] The user accesses the system via their device and logs in. Upon successful login, a text input form for content generation is displayed. The user enters the text that will be the source of the content they want to generate into the form and submits it.
[0484] Step 2:
[0485] The terminal forwards the transmitted text to the server. The server receives this text and performs tokenization, syntactic analysis, and semantic analysis of the text using a natural language processing (NLP) module. This analysis extracts the main points and context of the text.
[0486] Step 3:
[0487] Users can select the style and genre of visual and auditory content on their device. For example, options such as fantasy or scientific style, narration, and music are provided. Once selections are complete, the device sends that information to the server.
[0488] Step 4:
[0489] The server generates visual content using an image generation algorithm based on the parsed text, taking into account the selected style and genre. Sound waveforms and lighting effects are also applied in this step.
[0490] Step 5:
[0491] The server runs a speech synthesis engine to generate audio content based on text. Speech parameters such as tone, speed, and accent are adjusted according to the selected genre.
[0492] Step 6:
[0493] When a user requests multilingual content, the server uses a translation module to convert the text into the specified language. After translation, the visual and auditory content is generated again.
[0494] Step 7:
[0495] The generated content can be shared and collaboratively edited among users. The server generates a collaboration link and sends it to the user's device. Users use this link to begin collaborating with other users.
[0496] Step 8:
[0497] The final generated content is displayed on the device. Users can choose to save or download it. Following the user's instructions, the server stores the content in a database for later reuse.
[0498] (Example 1)
[0499] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0500] Currently, many digital content generation systems generate visual or auditory information based on text entered by users in natural language. However, these systems have limited multilingual support and real-time collaborative editing capabilities, failing to adequately meet the diverse needs of users. Furthermore, the inability to smoothly share and collaborate on content among different users presents inconveniences.
[0501] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0502] In this invention, the server includes means for analyzing the meaning of an input document using natural language processing technology, means for generating visual information in a selected format, means for generating auditory information in a selected category, means for translating into different languages and generating content again, and means for synchronizing the generated content with other devices and collaborating. This enables the creation of more multilingual and collaborative digital content.
[0503] "Natural language processing technology" refers to a set of methods and algorithms that enable computers to understand and analyze human language.
[0504] A "document" refers to text data entered by a user, which is the subject of natural language processing.
[0505] "Format" refers to the specific style or design pattern that a user chooses when generating visual information.
[0506] "Visual information" refers to content generated to create visual representations such as images and illustrations.
[0507] A "category" refers to a specific genre or theme that a user selects when generating auditory information.
[0508] "Auditory information" refers to content generated to create auditory expressions such as speech and music.
[0509] Translation is the process of converting a document from one language to another.
[0510] "Device" refers to hardware or software used in conjunction with a system to display and edit digital content.
[0511] "Synchronization" is the process of coordinating data across multiple devices to ensure consistency and updating information in real time.
[0512] "Collaborative work" refers to activities in which multiple users work together to edit or create a single piece of content, either simultaneously or at different times.
[0513] The system in this invention uses natural language processing technology to analyze documents input by a user and generate visual and auditory information. Specifically, a server, a terminal, and a user cooperate to carry out the digital content generation process.
[0514] The user begins the process by accessing the system using their device and logging in. After the login screen, the user is presented with a text input form. The user enters a prompt into this form, such as "Please create a story with the theme of the arrival of spring." The device then sends this input document to the server.
[0515] The server analyzes received documents using natural language processing (NLP) techniques. It performs text tokenization, syntactic analysis, and semantic analysis, and based on the results, generates visual and auditory information according to the format and category selected by the user. Visual information is created using image generation software, and auditory information is created using a speech synthesis engine.
[0516] Furthermore, the server uses translation modules as needed to convert content into other languages. This enables the generation of multilingual content. The generated content is sent from the server to the terminal in real time, and users can view the results on their terminal and download or share them.
[0517] Furthermore, the system includes collaborative features, allowing users to share generated content with other users in real time and collaboratively edit it. In this way, the invention transforms users' ideas into visual and auditory digital content, enabling efficient and flexible content creation.
[0518] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0519] Step 1:
[0520] The user accesses the system through their terminal and logs in. After logging in, a text input form is displayed. The user enters a prompt message into this form—for example, "Please create a story with the theme of the arrival of spring." The entered text is treated as the system's initial data.
[0521] Step 2:
[0522] The terminal sends the prompt message entered by the user to the server. At this time, the prompt message is encoded as a digital signal and transferred to the server according to the communication protocol. The server receives this signal, decodes it as text data, and uses it as the start data for processing.
[0523] Step 3:
[0524] The server analyzes the received text data using a natural language processing (NLP) module. Specifically, it tokenizes the input sentence and analyzes its grammatical structure using a syntactic analysis engine. Furthermore, it understands the context through semantic analysis and extracts the information necessary for generating content. The analysis results become the foundational data for subsequent data generation processes.
[0525] Step 4:
[0526] Based on the analysis results, the server generates visual and auditory information according to the format and category specified by the user. Visual information is created and output using an image generation algorithm. Auditory information is output as audio or music according to the specified audio parameters using a speech synthesis engine. Image files and audio files are generated during this process.
[0527] Step 5:
[0528] The server sends the generated visual and auditory information to the user's device. This output content is displayed on the device in real time. The user can review the generated information on their device and save or download it as needed.
[0529] Step 6:
[0530] If necessary, the user requests content generation in another language from their device. The server uses a translation module to convert existing text data into the other language. Then, it generates visual and auditory information again and sends it to the user's device. This completes the delivery of multilingual content.
[0531] Step 7:
[0532] Users can share generated information with other users within the system and collaboratively edit it. Content can be modified or added in real time via the terminal, the server reflects these changes, and the updated data is immediately sent to all participants. This enables efficient collaboration.
[0533] (Application Example 1)
[0534] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0535] In content distribution services, there is a need to provide an environment where users can easily generate and share diverse visual and auditory information, and furthermore, to realize a function that allows for real-time collaborative editing of works with other users. Existing systems have problems that hinder user convenience, such as requiring specialized technical skills for content generation and editing, and having insufficient multilingual support.
[0536] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0537] In this invention, the server includes means for analyzing input text using natural language processing technology, means for generating visual information in a specified format based on the analysis results, means for generating auditory information of a specified type based on the analysis results, and means for displaying the visual and auditory information in a virtual environment. This enables users to generate and share diverse visual and auditory information and to perform real-time collaborative editing that is compatible with other languages, even without specialized technical skills.
[0538] "Natural language processing technology" refers to the technology that enables computers to understand and process human language appropriately.
[0539] "Means of analysis" refers to a device or program that has the function of analyzing the characteristics and context of input text.
[0540] "Style" refers to the genre or style of visual information or works of art, and is associated with specific designs and methods of expression.
[0541] "Visual information" refers to images, illustrations, videos, etc., generated based on the analysis results.
[0542] "Type" refers to the category or genre of auditory information or musical works, and is associated with a specific musical style or sound profile.
[0543] "Auditory information" refers to speech, sounds, music, etc., generated based on the analysis results.
[0544] A "virtual environment" refers to a space for displaying and manipulating information on a virtual space or digital platform.
[0545] "Sharing" refers to a state in which multiple users can simultaneously view and access information or content.
[0546] "Collaborative editing" refers to a state in which multiple users can simultaneously edit and modify a single work or piece of information.
[0547] This invention provides a system that generates visual and auditory information based on text data entered by a user, and specific embodiments for implementing this system are shown below.
[0548] Users access this system using devices such as smartphones and personal computers. First, users enter text content into input forms displayed on their devices. The server then analyzes the text data using appropriate natural language processing technologies, such as the spaCy library. This analysis process includes tokenization, syntactic analysis, and semantic analysis to deepen the understanding of the language.
[0549] Based on the analyzed data, the server generates visual information in a specified format using image generation algorithms such as Stable Diffusion. It also generates auditory information of a specified type using a speech synthesis engine such as Google Text-to-Speech. The generated information is displayed in a virtual environment on a digital platform, which users can view and interact with.
[0550] Furthermore, anticipating that content will be used in multiple languages, a translation module for multilingual support (e.g., Google Translate API) has been implemented. This ensures that users can access content in different languages.
[0551] As a concrete example, if a user enters the beginning of "The Little Prince" and selects the "Fantasy" style, the server will analyze the entered text and generate a fantastical illustration and audio narration that matches the story. An example of a prompt in this case might be, "Please create fantasy-style illustrations and audio content from the text."
[0552] This system allows users to create and share engaging digital content without requiring specialized technical skills.
[0553] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0554] Step 1:
[0555] The user accesses the system using a terminal and logs in. The user enters the text they want to generate into a text input form on the terminal. This input data is sent to the server as basic information for subsequent analysis and generation.
[0556] Step 2:
[0557] The server receives text data sent by the user. Based on the received data, it starts analysis using natural language processing (NLP) techniques. Specifically, it tokenizes the text using an NLP library such as spaCy, and then performs syntactic and semantic analysis. This process captures the characteristics and context of the text, and outputs the basic data for the next generation step.
[0558] Step 3:
[0559] Users select the style and genre of content to be generated on their device. This selection is sent to the server and reflected in the generation process. This selection influences the algorithms used to generate visual and auditory content.
[0560] Step 4:
[0561] The server generates visual information using image generation algorithms such as Stable Diffusion, based on the analyzed text data and user selection information. In this step, the analyzed data and format are used as input, and the output is a graphic in the specified style. The generated visual information is temporarily stored on the server.
[0562] Step 5:
[0563] In parallel, the server uses a text-to-speech engine, such as Google Text-to-Speech, to generate auditory information based on the parsed text data. The parsed data and type are used as input, and the output is audio content of the specified type. This generates natural-sounding speech that reflects the user's intent.
[0564] Step 6:
[0565] The server generates content in multiple languages, using translation modules (e.g., Google Translate API) as needed. In this step, the original text and the user-selected language are used as input, and the output is the translated text and the corresponding audio content.
[0566] Step 7:
[0567] The generated visual and auditory information is sent to the user's device and displayed in the virtual environment. The user can view and interact with this information. Furthermore, the displayed content can be shared and collaboratively edited with other users, facilitating real-time communication and creative work.
[0568] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0569] The present invention is a visual and auditory digital content generation system that combines an emotion engine that recognizes user emotions. Embodiments of this system will be described below.
[0570] Overall structure and user interface
[0571] Users log in to the system via their device and enter the necessary information into a text input form. This entered text becomes the basic data for the generated digital content. Users can then select visual styles and auditory genres to enjoy a seamless experience.
[0572] Emotion recognition and data analysis
[0573] The server analyzes the text sent from the terminal using natural language processing techniques. During this analysis, an emotion engine recognizes the user's emotions and determines the emotional tone based on the text. The recognized emotions are then used to dynamically adjust the subsequent content generation process.
[0574] Content generation process
[0575] Based on the analysis results and the emotion engine's output, the server generates visual digital content that takes style selection and emotional tone into consideration. For example, if the user is excited, a colorful and dynamic design will be chosen. Similarly, auditory content is generated based on emotion, with the speech synthesis engine adjusting the tone and tempo.
[0576] Multilingual support and collaboration features
[0577] The server's multilingual translation feature provides users with the option to generate content in different languages. The generated content can be shared with other users in real time, and collaboration features facilitate co-editing. This feature is particularly useful when project teams are multinational.
[0578] Content output and saving
[0579] The generated visual and auditory content is displayed on the user's device. The user can review the content and choose to download it or save it on the server as needed. The server stores the content in a database accordingly and manages it as a history.
[0580] In this way, systems that combine emotion engines provide personalized digital content that reflects the user's emotional state, aiming to improve creativity and efficiency. For example, if a user inputs a part of an emotionally moving story, digital content is generated consisting of a visual design with calming colors and emotionally moving music.
[0581] The following describes the processing flow.
[0582] Step 1:
[0583] The user logs into the system using their device and accesses the dashboard. They enter the text that will be the source of the content they want to generate into the displayed text input form and click the "Submit" button.
[0584] Step 2:
[0585] The terminal sends the text data entered by the user to the emotion engine. The server receives this text and analyzes its content using natural language processing techniques.
[0586] Step 3:
[0587] The server extracts contextual information from the analyzed text and identifies the user's emotions through an emotion engine. For example, the emotion engine determines whether the text contains emotions such as joy, sadness, or excitement.
[0588] Step 4:
[0589] The user receives an option on their device to select their preferred visual style and auditory genre. For example, they might choose "animation style" or "classical music."
[0590] Step 5:
[0591] The server begins generating visual content based on the emotions recognized by the emotion engine and the user's selections. The server uses an image generation algorithm to dynamically determine colors and designs that correspond to the emotions.
[0592] Step 6:
[0593] Similarly, the server operates a speech synthesis engine to generate auditory content. Here, the tone and tempo of the voice are adjusted to match the user's emotions, creating a sound that fits those emotions.
[0594] Step 7:
[0595] If the user has selected multilingual support, the server will use a translation API to convert the text into other languages and regenerate the content as needed.
[0596] Step 8:
[0597] The generated visual and auditory content is displayed on the device. Users can review the results and, if satisfied, download the content or share it with others using collaboration tools.
[0598] Step 9:
[0599] If a user chooses to save content, the server stores it in a database and manages it as part of the user's project history. This allows for later reuse and editing.
[0600] (Example 2)
[0601] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0602] A challenge in modern digital content creation is the difficulty in generating content that fully reflects the user's emotions and intentions. Conventional generation methods are limited to simple conversions based on input text data, making it difficult to provide customized content tailored to the specific emotions or situations expressed by the user. Furthermore, support for content sharing and collaborative editing across different languages is insufficient.
[0603] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0604] In this invention, the server includes means for performing sentiment analysis to recognize the user's emotions, means for dynamically adjusting and generating visual and auditory digital content based on the emotions, and means for generating customized content that reflects the user's input information and emotions. This enables the provision of personalized digital content that enhances the user experience.
[0605] A "user" is a person or group that operates the system and provides input data.
[0606] "Sentiment analysis" is the process of detecting a user's emotional state from input text data and interpreting the data based on that.
[0607] "Visual digital content" refers to content that includes images, videos, or visual elements generated in a digital format.
[0608] "Auditory digital content" refers to content that generates sound-related elements, such as voices and music, in a digital format.
[0609] "Natural language processing" is the technology that enables computers to understand, interpret, and manipulate human language.
[0610] "Multilingual translation" is the process of translating data between different languages and providing the translated result in a specific language.
[0611] "Collaborative work tools" are features that allow multiple users to jointly edit or share content.
[0612] A "database" is a system for organizing, efficiently storing, searching, and managing data.
[0613] This system is designed to allow users to generate digital content via their devices. Users log in using an interface on their devices and then input text data. This text data is sent to the server and used as the basis for the digital content.
[0614] The server uses NLP libraries and frameworks (e.g., SpaCy, NLTK) to analyze this text data using natural language processing techniques. Based on this analysis, it runs an emotion analysis engine to recognize the user's emotions. The server then uses generative AI models (e.g., DALL-E, VQ-VAE-2) to generate visual and auditory digital content. During this process, a connected speech synthesis engine adjusts the tone and tempo of the generated audio data.
[0615] The server uses a multilingual translation API to translate generated content into different languages, enabling the provision of content that supports various language environments. Furthermore, the generated content is designed to be shared and collaboratively edited with other users in real time. For this purpose, it incorporates collaborative work tools, allowing for smooth operation even when project teams are composed of members from multiple countries.
[0616] The content is displayed on the user's device, and after viewing it, the user can download or save it to the server as needed. The server manages the content in a database and stores it as a history so that the user can access it later.
[0617] For example, if a user enters the prompt "Generate a heartwarming story," the server will generate digital content including a visual design with gentle, calming colors and moving music. In this way, it becomes possible to provide personalized digital content based on the user's intentions and emotions.
[0618] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0619] Step 1:
[0620] Users log in to the system via their device. During login, the user enters authentication information, which the device then sends to the server. The server verifies this information to confirm that the user is a legitimate user. If the login is successful, the user can access the system.
[0621] Step 2:
[0622] The user enters text data on the terminal interface. In the input form, the user provides information related to the digital content in text format. The terminal collects this input data and sends it to the server. This entered text data is used as the basis for the digital content to be generated.
[0623] Step 3:
[0624] The server performs natural language processing on the received text data. The specific software used is an NLP library such as SpaCy or NLTK, which performs grammatical and semantic analysis. The input data is analyzed, and as a result, the text's structure and semantic information are extracted. This information obtained from the analysis is then used for subsequent processing.
[0625] Step 4:
[0626] The server uses the analysis results to perform sentiment analysis. This process extracts the emotions contained in the input text using a sentiment analysis engine. The analysis results are then used during subsequent content generation, forming the basis for data adjustments based on the user's emotional state.
[0627] Step 5:
[0628] The server uses generative AI models to generate visual and auditory digital content. Specifically, it uses generative AI models such as DALL-E and VQ-VAE-2 to create content that matches the user's emotions. In this process, the design's color scheme and audio tone are set based on the emotional tone. The output content is then provided to the user.
[0629] Step 6:
[0630] The server uses a multilingual translation API to translate the generated content into other languages. This makes the content understandable to users in multilingual environments. The server processes the translated data in real time and generates output that includes corresponding language information.
[0631] Step 7:
[0632] The server sends the generated content to the user's device and simultaneously generates a link for sharing with other users. Users can use this link in real time to collaborate on editing with other users. The generated link information is passed around via email, chat, etc.
[0633] Step 8:
[0634] The user reviews the content received on their device and downloads or saves it to the server as needed. The server, following the user's instructions, stores the content in a database and manages it as a history. This makes it easy for the user to access the content at a later date.
[0635] (Application Example 2)
[0636] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0637] In the modern era, digital content is delivered in a variety of formats, but there is a challenge in the lack of dynamic content adjustments that respond to user emotions. Furthermore, there is a demand for content that is more personalized and relatable to viewers. To achieve this, technology is needed that accurately analyzes user emotions and dynamically generates and adjusts content based on those results.
[0638] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0639] In this invention, the server includes means for analyzing the meaning of input text using natural language processing technology, means for analyzing the user's emotions using emotion recognition technology and dynamically adjusting visual and auditory digital content according to those emotions, and means for changing the tone and theme of a video according to the user's emotions. This makes it possible to provide personalized digital content based on the user's emotions.
[0640] "Natural language processing technology" is a technology that uses computers to understand and analyze human language.
[0641] "Text semantic analysis" involves structurally analyzing input text to understand its content and intent.
[0642] "Generating visual digital content" means creating media that can be displayed in a digital format based on visual information.
[0643] "Generating auditory digital content" means recording or generating auditory information in a digital format to create playable audio data.
[0644] "Emotion recognition technology" is a technology that detects and analyzes the emotions expressed by a user as data.
[0645] "Dynamically adjusting content" means changing the nature or structure of content in real time or as needed.
[0646] "Changing the tone or theme of a video" means altering the atmosphere or message of the video to provide viewers with a different visual experience.
[0647] This invention is a system that allows users to access the system via a terminal to generate personalized visual and auditory digital content based on their emotions. Specific embodiments thereof are described below.
[0648] The server first analyzes the text data received from the user using natural language processing techniques. This analysis extracts the user's intentions and emotions from the text. Specifically, Python's natural language processing library (e.g., NLTK) may be used for this analysis.
[0649] Based on the analysis results, the server then uses emotion recognition technology to understand the user's emotions in detail. In this process, emotion recognition technology identifies the user's emotions and determines what tone and theme of content is appropriate.
[0650] Based on the emotional information received, the server generates visual and auditory digital content. The visual content is dynamically styled, with changes in color tone and brightness as needed. The auditory content, on the other hand, uses a speech synthesis engine to apply emotionally appropriate tone and tempo.
[0651] For example, a prompt message given to the system might be, "Design the best visual content and music for when the user wants to relax." Based on this instruction, the system generates and provides a video of a calm nighttime landscape and soothing music.
[0652] Finally, the generated content is sent to the user's device and displayed or played. At this point, the user can review the generated content and determine if it is appropriate. Furthermore, the content can be generated in multiple languages and can be shared and collaboratively edited among users. The server records the generation process and stores it as data for future content generation.
[0653] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0654] Step 1:
[0655] The terminal receives text data from the user as input. The user uses prompts to specify the intent and content of the material they want to generate. This input text data is sent to the server as initial system data.
[0656] Step 2:
[0657] The server analyzes the received text data using natural language processing techniques. Specifically, it extracts nouns, verbs, and other elements from the input text using a natural language processing library to understand the context and sentiment. The output of this analysis contains information about the user's intent and sentiment, and serves as foundational data for the next step.
[0658] Step 3:
[0659] The server applies emotion recognition technology based on the analysis results to identify the user's emotions. Based on the analysis data received as input, the emotion engine determines the intensity and type of emotion. For example, emotions such as joy and sadness are extracted, and the results influence subsequent content generation.
[0660] Step 4:
[0661] The server takes identified emotional information as input and generates visual and auditory digital content that matches the emotion. Visual content is dynamically adjusted in terms of color and layout, while auditory content has its tone and tempo set using a speech synthesis engine. The output is customized digital content that matches the user's emotion.
[0662] Step 5:
[0663] The server sends the generated content to the terminal. The user can experience the output visually or aurally and judge the appropriateness of the content.
[0664] Step 6:
[0665] The server records the generated content and its generation process, and stores it in a database. This information is used to improve the content generation process in the future and to understand user preferences.
[0666] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0667] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0668] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0669] [Fourth Embodiment]
[0670] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0671] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0672] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0673] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0674] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0675] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0676] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0677] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0678] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0679] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0680] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0681] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0682] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0683] This invention provides users with a system for generating visual and auditory digital content using natural language processing technology. The following is an example of this system.
[0684] User interface and text input
[0685] The user accesses the system using their device and logs in to begin the content generation process. A text input form appears on the device, and the user enters text related to the content they want to generate into this form. For example, this could be a portion of a novel or the beginning of a blog post.
[0686] Data reception and analysis
[0687] The server receives text data sent from the user's terminal. The received text is parsed by the server's natural language processing (NLP) module. Specifically, it performs tokenization, syntactic analysis, and semantic analysis of the text to determine what visual and auditory content is appropriate based on its characteristics and context.
[0688] Style and Genre Selection
[0689] Users can also select the style and genre of the generated content on their device. This selection allows them to choose the most appropriate style or genre from several options as needed, and is an important step in reflecting the user's intentions.
[0690] Content generation
[0691] The server uses an image generation algorithm to create visual content based on the analyzed text data and the style and genre selected by the user. Specific illustrations and graphics are created during this process. Similarly, a speech synthesis engine is used to generate auditory content. At this time, speech parameters are set according to the user's selection to produce natural-sounding speech and music.
[0692] Multilingual support and collaboration
[0693] Upon user request, the server can use a translation module to convert text into another language and regenerate the content. Furthermore, collaboration features are provided for sharing and co-editing the generated content with other users in real time. This allows users to efficiently advance their projects.
[0694] Output and save
[0695] The final generated visual and auditory content is displayed on the user's device, and the user can choose to download or save it. The content can be customized to the user's needs and is available across various digital media.
[0696] This system allows users to easily generate diverse digital content without requiring advanced professional skills. This significantly contributes to improving the efficiency and creativity of content creation.
[0697] The following describes the processing flow.
[0698] Step 1:
[0699] The user accesses the system via their device and logs in. Upon successful login, a text input form for content generation is displayed. The user enters the text that will be the source of the content they want to generate into the form and submits it.
[0700] Step 2:
[0701] The terminal forwards the transmitted text to the server. The server receives this text and performs tokenization, syntactic analysis, and semantic analysis of the text using a natural language processing (NLP) module. This analysis extracts the main points and context of the text.
[0702] Step 3:
[0703] Users can select the style and genre of visual and auditory content on their device. For example, options such as fantasy or scientific style, narration, and music are provided. Once selections are complete, the device sends that information to the server.
[0704] Step 4:
[0705] The server generates visual content using an image generation algorithm based on the parsed text, taking into account the selected style and genre. Sound waveforms and lighting effects are also applied in this step.
[0706] Step 5:
[0707] The server runs a speech synthesis engine to generate audio content based on text. Speech parameters such as tone, speed, and accent are adjusted according to the selected genre.
[0708] Step 6:
[0709] When a user requests multilingual content, the server uses a translation module to convert the text into the specified language. After translation, the visual and auditory content is generated again.
[0710] Step 7:
[0711] The generated content can be shared and collaboratively edited among users. The server generates a collaboration link and sends it to the user's device. Users use this link to begin collaborating with other users.
[0712] Step 8:
[0713] The final generated content is displayed on the device. Users can choose to save or download it. Following the user's instructions, the server stores the content in a database for later reuse.
[0714] (Example 1)
[0715] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0716] Currently, many digital content generation systems generate visual or auditory information based on text entered by users in natural language. However, these systems have limited multilingual support and real-time collaborative editing capabilities, failing to adequately meet the diverse needs of users. Furthermore, the inability to smoothly share and collaborate on content among different users presents inconveniences.
[0717] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0718] In this invention, the server includes means for analyzing the meaning of an input document using natural language processing technology, means for generating visual information in a selected format, means for generating auditory information in a selected category, means for translating into different languages and generating content again, and means for synchronizing the generated content with other devices and collaborating. This enables the creation of more multilingual and collaborative digital content.
[0719] "Natural language processing technology" refers to a set of methods and algorithms that enable computers to understand and analyze human language.
[0720] A "document" refers to text data entered by a user, which is the subject of natural language processing.
[0721] "Format" refers to the specific style or design pattern that a user chooses when generating visual information.
[0722] "Visual information" refers to content generated to create visual representations such as images and illustrations.
[0723] A "category" refers to a specific genre or theme that a user selects when generating auditory information.
[0724] "Auditory information" refers to content generated to create auditory expressions such as speech and music.
[0725] Translation is the process of converting a document from one language to another.
[0726] "Device" refers to hardware or software used in conjunction with a system to display and edit digital content.
[0727] "Synchronization" is the process of coordinating data across multiple devices to ensure consistency and updating information in real time.
[0728] "Collaborative work" refers to activities in which multiple users work together to edit or create a single piece of content, either simultaneously or at different times.
[0729] The system in this invention uses natural language processing technology to analyze documents input by a user and generate visual and auditory information. Specifically, a server, a terminal, and a user cooperate to carry out the digital content generation process.
[0730] The user begins the process by accessing the system using their device and logging in. After the login screen, the user is presented with a text input form. The user enters a prompt into this form, such as "Please create a story with the theme of the arrival of spring." The device then sends this input document to the server.
[0731] The server analyzes received documents using natural language processing (NLP) techniques. It performs text tokenization, syntactic analysis, and semantic analysis, and based on the results, generates visual and auditory information according to the format and category selected by the user. Visual information is created using image generation software, and auditory information is created using a speech synthesis engine.
[0732] Furthermore, the server uses translation modules as needed to convert content into other languages. This enables the generation of multilingual content. The generated content is sent from the server to the terminal in real time, and users can view the results on their terminal and download or share them.
[0733] Furthermore, the system includes collaborative features, allowing users to share generated content with other users in real time and collaboratively edit it. In this way, the invention transforms users' ideas into visual and auditory digital content, enabling efficient and flexible content creation.
[0734] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0735] Step 1:
[0736] The user accesses the system through their terminal and logs in. After logging in, a text input form is displayed. The user enters a prompt message into this form—for example, "Please create a story with the theme of the arrival of spring." The entered text is treated as the system's initial data.
[0737] Step 2:
[0738] The terminal sends the prompt message entered by the user to the server. At this time, the prompt message is encoded as a digital signal and transferred to the server according to the communication protocol. The server receives this signal, decodes it as text data, and uses it as the start data for processing.
[0739] Step 3:
[0740] The server analyzes the received text data using a natural language processing (NLP) module. Specifically, it tokenizes the input sentence and analyzes its grammatical structure using a syntactic analysis engine. Furthermore, it understands the context through semantic analysis and extracts the information necessary for generating content. The analysis results become the foundational data for subsequent data generation processes.
[0741] Step 4:
[0742] Based on the analysis results, the server generates visual and auditory information according to the format and category specified by the user. Visual information is created and output using an image generation algorithm. Auditory information is output as audio or music according to the specified audio parameters using a speech synthesis engine. Image files and audio files are generated during this process.
[0743] Step 5:
[0744] The server sends the generated visual and auditory information to the user's device. This output content is displayed on the device in real time. The user can review the generated information on their device and save or download it as needed.
[0745] Step 6:
[0746] If necessary, the user requests content generation in another language from their device. The server uses a translation module to convert existing text data into the other language. Then, it generates visual and auditory information again and sends it to the user's device. This completes the delivery of multilingual content.
[0747] Step 7:
[0748] Users can share generated information with other users within the system and collaboratively edit it. Content can be modified or added in real time via the terminal, the server reflects these changes, and the updated data is immediately sent to all participants. This enables efficient collaboration.
[0749] (Application Example 1)
[0750] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0751] In content distribution services, there is a need to provide an environment where users can easily generate and share diverse visual and auditory information, and furthermore, to realize a function that allows for real-time collaborative editing of works with other users. Existing systems have problems that hinder user convenience, such as requiring specialized technical skills for content generation and editing, and having insufficient multilingual support.
[0752] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0753] In this invention, the server includes means for analyzing input text using natural language processing technology, means for generating visual information in a specified format based on the analysis results, means for generating auditory information of a specified type based on the analysis results, and means for displaying the visual and auditory information in a virtual environment. This enables users to generate and share diverse visual and auditory information and to perform real-time collaborative editing that is compatible with other languages, even without specialized technical skills.
[0754] "Natural language processing technology" refers to the technology that enables computers to understand and process human language appropriately.
[0755] "Means of analysis" refers to a device or program that has the function of analyzing the characteristics and context of input text.
[0756] "Style" refers to the genre or style of visual information or works of art, and is associated with specific designs and methods of expression.
[0757] "Visual information" refers to images, illustrations, videos, etc., generated based on the analysis results.
[0758] "Type" refers to the category or genre of auditory information or musical works, and is associated with a specific musical style or sound profile.
[0759] "Auditory information" refers to speech, sounds, music, etc., generated based on the analysis results.
[0760] A "virtual environment" refers to a space for displaying and manipulating information on a virtual space or digital platform.
[0761] "Sharing" refers to a state in which multiple users can simultaneously view and access information or content.
[0762] "Collaborative editing" refers to a state in which multiple users can simultaneously edit and modify a single work or piece of information.
[0763] This invention provides a system that generates visual and auditory information based on text data entered by a user, and specific embodiments for implementing this system are shown below.
[0764] Users access this system using devices such as smartphones and personal computers. First, users enter text content into input forms displayed on their devices. The server then analyzes the text data using appropriate natural language processing technologies, such as the spaCy library. This analysis process includes tokenization, syntactic analysis, and semantic analysis to deepen the understanding of the language.
[0765] Based on the analyzed data, the server generates visual information in a specified format using image generation algorithms such as Stable Diffusion. It also generates auditory information of a specified type using a speech synthesis engine such as Google Text-to-Speech. The generated information is displayed in a virtual environment on a digital platform, which users can view and interact with.
[0766] Furthermore, anticipating that content will be used in multiple languages, a translation module for multilingual support (e.g., Google Translate API) has been implemented. This ensures that users can access content in different languages.
[0767] As a concrete example, if a user enters the beginning of "The Little Prince" and selects the "Fantasy" style, the server will analyze the entered text and generate a fantastical illustration and audio narration that matches the story. An example of a prompt in this case might be, "Please create fantasy-style illustrations and audio content from the text."
[0768] This system allows users to create and share engaging digital content without requiring specialized technical skills.
[0769] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0770] Step 1:
[0771] The user accesses the system using a terminal and logs in. The user enters the text they want to generate into a text input form on the terminal. This input data is sent to the server as basic information for subsequent analysis and generation.
[0772] Step 2:
[0773] The server receives text data sent by the user. Based on the received data, it starts analysis using natural language processing (NLP) techniques. Specifically, it tokenizes the text using an NLP library such as spaCy, and then performs syntactic and semantic analysis. This process captures the characteristics and context of the text, and outputs the basic data for the next generation step.
[0774] Step 3:
[0775] Users select the style and genre of content to be generated on their device. This selection is sent to the server and reflected in the generation process. This selection influences the algorithms used to generate visual and auditory content.
[0776] Step 4:
[0777] The server generates visual information using image generation algorithms such as Stable Diffusion, based on the analyzed text data and user selection information. In this step, the analyzed data and format are used as input, and the output is a graphic in the specified style. The generated visual information is temporarily stored on the server.
[0778] Step 5:
[0779] In parallel, the server uses a text-to-speech engine, such as Google Text-to-Speech, to generate auditory information based on the parsed text data. The parsed data and type are used as input, and the output is audio content of the specified type. This generates natural-sounding speech that reflects the user's intent.
[0780] Step 6:
[0781] The server generates content in multiple languages, using translation modules (e.g., Google Translate API) as needed. In this step, the original text and the user-selected language are used as input, and the output is the translated text and the corresponding audio content.
[0782] Step 7:
[0783] The generated visual and auditory information is sent to the user's device and displayed in the virtual environment. The user can view and interact with this information. Furthermore, the displayed content can be shared and collaboratively edited with other users, facilitating real-time communication and creative work.
[0784] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0785] The present invention is a visual and auditory digital content generation system that combines an emotion engine that recognizes user emotions. Embodiments of this system will be described below.
[0786] Overall structure and user interface
[0787] Users log in to the system via their device and enter the necessary information into a text input form. This entered text becomes the basic data for the generated digital content. Users can then select visual styles and auditory genres to enjoy a seamless experience.
[0788] Emotion recognition and data analysis
[0789] The server analyzes the text sent from the terminal using natural language processing techniques. During this analysis, an emotion engine recognizes the user's emotions and determines the emotional tone based on the text. The recognized emotions are then used to dynamically adjust the subsequent content generation process.
[0790] Content generation process
[0791] Based on the analysis results and the emotion engine's output, the server generates visual digital content that takes style selection and emotional tone into consideration. For example, if the user is excited, a colorful and dynamic design will be chosen. Similarly, auditory content is generated based on emotion, with the speech synthesis engine adjusting the tone and tempo.
[0792] Multilingual support and collaboration features
[0793] The server's multilingual translation feature provides users with the option to generate content in different languages. The generated content can be shared with other users in real time, and collaboration features facilitate co-editing. This feature is particularly useful when project teams are multinational.
[0794] Content output and saving
[0795] The generated visual and auditory content is displayed on the user's device. The user can review the content and choose to download it or save it on the server as needed. The server stores the content in a database accordingly and manages it as a history.
[0796] In this way, systems that combine emotion engines provide personalized digital content that reflects the user's emotional state, aiming to improve creativity and efficiency. For example, if a user inputs a part of an emotionally moving story, digital content is generated consisting of a visual design with calming colors and emotionally moving music.
[0797] The following describes the processing flow.
[0798] Step 1:
[0799] The user logs into the system using their device and accesses the dashboard. They enter the text that will be the source of the content they want to generate into the displayed text input form and click the "Submit" button.
[0800] Step 2:
[0801] The terminal sends the text data entered by the user to the emotion engine. The server receives this text and analyzes its content using natural language processing techniques.
[0802] Step 3:
[0803] The server extracts contextual information from the analyzed text and identifies the user's emotions through an emotion engine. For example, the emotion engine determines whether the text contains emotions such as joy, sadness, or excitement.
[0804] Step 4:
[0805] The user receives an option on their device to select their preferred visual style and auditory genre. For example, they might choose "animation style" or "classical music."
[0806] Step 5:
[0807] The server begins generating visual content based on the emotions recognized by the emotion engine and the user's selections. The server uses an image generation algorithm to dynamically determine colors and designs that correspond to the emotions.
[0808] Step 6:
[0809] Similarly, the server operates a speech synthesis engine to generate auditory content. Here, the tone and tempo of the voice are adjusted to match the user's emotions, creating a sound that fits those emotions.
[0810] Step 7:
[0811] If the user has selected multilingual support, the server will use a translation API to convert the text into other languages and regenerate the content as needed.
[0812] Step 8:
[0813] The generated visual and auditory content is displayed on the device. Users can review the results and, if satisfied, download the content or share it with others using collaboration tools.
[0814] Step 9:
[0815] If a user chooses to save content, the server stores it in a database and manages it as part of the user's project history. This allows for later reuse and editing.
[0816] (Example 2)
[0817] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0818] A challenge in modern digital content creation is the difficulty in generating content that fully reflects the user's emotions and intentions. Conventional generation methods are limited to simple conversions based on input text data, making it difficult to provide customized content tailored to the specific emotions or situations expressed by the user. Furthermore, support for content sharing and collaborative editing across different languages is insufficient.
[0819] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0820] In this invention, the server includes means for performing sentiment analysis to recognize the user's emotions, means for dynamically adjusting and generating visual and auditory digital content based on the emotions, and means for generating customized content that reflects the user's input information and emotions. This enables the provision of personalized digital content that enhances the user experience.
[0821] A "user" is a person or group that operates the system and provides input data.
[0822] "Sentiment analysis" is the process of detecting a user's emotional state from input text data and interpreting the data based on that.
[0823] "Visual digital content" refers to content that includes images, videos, or visual elements generated in a digital format.
[0824] "Auditory digital content" refers to content that generates sound-related elements, such as voices and music, in a digital format.
[0825] "Natural language processing" is the technology that enables computers to understand, interpret, and manipulate human language.
[0826] "Multilingual translation" is the process of translating data between different languages and providing the translated result in a specific language.
[0827] "Collaborative work tools" are features that allow multiple users to jointly edit or share content.
[0828] A "database" is a system for organizing, efficiently storing, searching, and managing data.
[0829] This system is designed to allow users to generate digital content via their devices. Users log in using an interface on their devices and then input text data. This text data is sent to the server and used as the basis for the digital content.
[0830] The server uses NLP libraries and frameworks (e.g., SpaCy, NLTK) to analyze this text data using natural language processing techniques. Based on this analysis, it runs an emotion analysis engine to recognize the user's emotions. The server then uses generative AI models (e.g., DALL-E, VQ-VAE-2) to generate visual and auditory digital content. During this process, a connected speech synthesis engine adjusts the tone and tempo of the generated audio data.
[0831] The server uses a multilingual translation API to translate generated content into different languages, enabling the provision of content that supports various language environments. Furthermore, the generated content is designed to be shared and collaboratively edited with other users in real time. For this purpose, it incorporates collaborative work tools, allowing for smooth operation even when project teams are composed of members from multiple countries.
[0832] The content is displayed on the user's device, and after viewing it, the user can download or save it to the server as needed. The server manages the content in a database and stores it as a history so that the user can access it later.
[0833] For example, if a user enters the prompt "Generate a heartwarming story," the server will generate digital content including a visual design with gentle, calming colors and moving music. In this way, it becomes possible to provide personalized digital content based on the user's intentions and emotions.
[0834] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0835] Step 1:
[0836] Users log in to the system via their device. During login, the user enters authentication information, which the device then sends to the server. The server verifies this information to confirm that the user is a legitimate user. If the login is successful, the user can access the system.
[0837] Step 2:
[0838] The user enters text data on the terminal interface. In the input form, the user provides information related to the digital content in text format. The terminal collects this input data and sends it to the server. This entered text data is used as the basis for the digital content to be generated.
[0839] Step 3:
[0840] The server performs natural language processing on the received text data. The specific software used is an NLP library such as SpaCy or NLTK, which performs grammatical and semantic analysis. The input data is analyzed, and as a result, the text's structure and semantic information are extracted. This information obtained from the analysis is then used for subsequent processing.
[0841] Step 4:
[0842] The server uses the analysis results to perform sentiment analysis. This process extracts the emotions contained in the input text using a sentiment analysis engine. The analysis results are then used during subsequent content generation, forming the basis for data adjustments based on the user's emotional state.
[0843] Step 5:
[0844] The server uses generative AI models to generate visual and auditory digital content. Specifically, it uses generative AI models such as DALL-E and VQ-VAE-2 to create content that matches the user's emotions. In this process, the design's color scheme and audio tone are set based on the emotional tone. The output content is then provided to the user.
[0845] Step 6:
[0846] The server uses a multilingual translation API to translate the generated content into other languages. This makes the content understandable to users in multilingual environments. The server processes the translated data in real time and generates output that includes corresponding language information.
[0847] Step 7:
[0848] The server sends the generated content to the user's device and simultaneously generates a link for sharing with other users. Users can use this link in real time to collaborate on editing with other users. The generated link information is passed around via email, chat, etc.
[0849] Step 8:
[0850] The user reviews the content received on their device and downloads or saves it to the server as needed. The server, following the user's instructions, stores the content in a database and manages it as a history. This makes it easy for the user to access the content at a later date.
[0851] (Application Example 2)
[0852] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0853] In the modern era, digital content is delivered in a variety of formats, but there is a challenge in the lack of dynamic content adjustments that respond to user emotions. Furthermore, there is a demand for content that is more personalized and relatable to viewers. To achieve this, technology is needed that accurately analyzes user emotions and dynamically generates and adjusts content based on those results.
[0854] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0855] In this invention, the server includes means for analyzing the meaning of input text using natural language processing technology, means for analyzing the user's emotions using emotion recognition technology and dynamically adjusting visual and auditory digital content according to those emotions, and means for changing the tone and theme of a video according to the user's emotions. This makes it possible to provide personalized digital content based on the user's emotions.
[0856] "Natural language processing technology" is a technology that uses computers to understand and analyze human language.
[0857] "Text semantic analysis" involves structurally analyzing input text to understand its content and intent.
[0858] "Generating visual digital content" means creating media that can be displayed in a digital format based on visual information.
[0859] "Generating auditory digital content" means recording or generating auditory information in a digital format to create playable audio data.
[0860] "Emotion recognition technology" is a technology that detects and analyzes the emotions expressed by a user as data.
[0861] "Dynamically adjusting content" means changing the nature or structure of content in real time or as needed.
[0862] "Changing the tone or theme of a video" means altering the atmosphere or message of the video to provide viewers with a different visual experience.
[0863] This invention is a system that allows users to access the system via a terminal to generate personalized visual and auditory digital content based on their emotions. Specific embodiments thereof are described below.
[0864] The server first analyzes the text data received from the user using natural language processing techniques. This analysis extracts the user's intentions and emotions from the text. Specifically, Python's natural language processing library (e.g., NLTK) may be used for this analysis.
[0865] Based on the analysis results, the server then uses emotion recognition technology to understand the user's emotions in detail. In this process, emotion recognition technology identifies the user's emotions and determines what tone and theme of content is appropriate.
[0866] Based on the emotional information received, the server generates visual and auditory digital content. The visual content is dynamically styled, with changes in color tone and brightness as needed. The auditory content, on the other hand, uses a speech synthesis engine to apply emotionally appropriate tone and tempo.
[0867] For example, a prompt message given to the system might be, "Design the best visual content and music for when the user wants to relax." Based on this instruction, the system generates and provides a video of a calm nighttime landscape and soothing music.
[0868] Finally, the generated content is sent to the user's device and displayed or played. At this point, the user can review the generated content and determine if it is appropriate. Furthermore, the content can be generated in multiple languages and can be shared and collaboratively edited among users. The server records the generation process and stores it as data for future content generation.
[0869] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0870] Step 1:
[0871] The terminal receives text data from the user as input. The user uses prompts to specify the intent and content of the material they want to generate. This input text data is sent to the server as initial system data.
[0872] Step 2:
[0873] The server analyzes the received text data using natural language processing techniques. Specifically, it extracts nouns, verbs, and other elements from the input text using a natural language processing library to understand the context and sentiment. The output of this analysis contains information about the user's intent and sentiment, and serves as foundational data for the next step.
[0874] Step 3:
[0875] The server applies emotion recognition technology based on the analysis results to identify the user's emotions. Based on the analysis data received as input, the emotion engine determines the intensity and type of emotion. For example, emotions such as joy and sadness are extracted, and the results influence subsequent content generation.
[0876] Step 4:
[0877] The server takes identified emotional information as input and generates visual and auditory digital content that matches the emotion. Visual content is dynamically adjusted in terms of color and layout, while auditory content has its tone and tempo set using a speech synthesis engine. The output is customized digital content that matches the user's emotion.
[0878] Step 5:
[0879] The server sends the generated content to the terminal. The user can experience the output visually or aurally and judge the appropriateness of the content.
[0880] Step 6:
[0881] The server records the generated content and its generation process, and stores it in a database. This information is used to improve the content generation process in the future and to understand user preferences.
[0882] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0883] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0884] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0885] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0886] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0887] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0888] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0889] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0890] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0891] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0892] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0893] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0894] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0895] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0896] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0897] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0898] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0899] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0900] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0901] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0902] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0903] The following is further disclosed regarding the embodiments described above.
[0904] (Claim 1)
[0905] A means of analyzing the meaning of input text using natural language processing technology,
[0906] A means for generating visual digital content in a specified style based on the results of the analysis of the aforementioned text,
[0907] A means for generating audio digital content in a specified genre based on the results of the analysis of the aforementioned text,
[0908] A system that includes this.
[0909] (Claim 2)
[0910] The system according to claim 1, further comprising means for generating the aforementioned visual digital content and auditory digital content in multiple languages.
[0911] (Claim 3)
[0912] The system according to claim 1, further comprising a collaboration means for enabling users to share and co-edit the aforementioned visual digital content and auditory digital content.
[0913] "Example 1"
[0914] (Claim 1)
[0915] A means of analyzing the meaning of an input document using natural language processing technology,
[0916] A means for generating visual information in a selected format based on the analysis results of the aforementioned document,
[0917] A means for generating auditory information in a selected category based on the analysis results of the aforementioned document,
[0918] Based on the aforementioned analysis results, a means for translating into different languages and regenerating the content,
[0919] A means of synchronizing the generated content with other devices and performing collaborative work,
[0920] A system that includes this.
[0921] (Claim 2)
[0922] The system according to claim 1, further comprising means for providing the results of generating the visual and auditory information in multiple languages.
[0923] (Claim 3)
[0924] The system according to claim 1, further comprising collaborative means for enabling users to share and jointly edit the results of generating the aforementioned visual and auditory information.
[0925] "Application Example 1"
[0926] (Claim 1)
[0927] A means of analyzing the meaning of input text using natural language processing technology,
[0928] A means for generating visual information in a specified format based on the results of the analysis of the aforementioned text,
[0929] A means for generating auditory information of a specified type based on the results of the analysis of the aforementioned text,
[0930] means for displaying the aforementioned visual and auditory information in a virtual environment,
[0931] A system that includes this.
[0932] (Claim 2)
[0933] The system according to claim 1, further comprising means for generating the aforementioned visual and auditory information in another language.
[0934] (Claim 3)
[0935] The system according to claim 1, further comprising means for enabling users to share and co-edit the aforementioned visual and auditory information.
[0936] "Example 2 of combining an emotion engine"
[0937] (Claim 1)
[0938] A means of performing sentiment analysis to recognize the user's emotions,
[0939] Means for dynamically adjusting and generating visual and auditory digital content based on the aforementioned emotions,
[0940] A means of generating customized content that reflects user input and emotions,
[0941] A natural language processing means for analyzing input text data,
[0942] A means for setting the style and genre of content based on the aforementioned analysis results and sentiment analysis results,
[0943] A means of performing real-time multilingual translation,
[0944] A system that includes this.
[0945] (Claim 2)
[0946] The system according to claim 1, further comprising collaborative means for sharing and jointly editing the aforementioned visual digital content and auditory digital content with other users.
[0947] (Claim 3)
[0948] The system according to claim 1, further comprising output means for displaying generated content to a user, and means for storing and managing the content in a database.
[0949] "Application example 2 of combining emotional engines"
[0950] (Claim 1)
[0951] A means of analyzing the meaning of input text using natural language processing technology,
[0952] A means for generating visual digital content in a specified style based on the results of the analysis of the aforementioned text,
[0953] A means for generating audio digital content in a specified genre based on the results of the analysis of the aforementioned text,
[0954] A means for analyzing a user's emotions using emotion recognition technology and dynamically adjusting visual and auditory digital content according to those emotions,
[0955] A means to change the tone and theme of a video according to the user's emotions,
[0956] A system that includes this.
[0957] (Claim 2)
[0958] The system according to claim 1, further comprising means for generating the aforementioned visual digital content and auditory digital content in multiple languages.
[0959] (Claim 3)
[0960] The system according to claim 1, further comprising a collaboration means for enabling users to share and co-edit the aforementioned visual digital content and auditory digital content. [Explanation of symbols]
[0961] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of analyzing the meaning of input text using natural language processing technology, A means for generating visual digital content in a specified style based on the results of the analysis of the aforementioned text, A means for generating audio digital content in a specified genre based on the results of the analysis of the aforementioned text, A system that includes this.
2. The system according to claim 1, further comprising means for generating the aforementioned visual digital content and auditory digital content in multiple languages.
3. The system according to claim 1, further comprising collaboration means for enabling users to share and co-edit the aforementioned visual digital content and auditory digital content.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A