system

A generative AI model optimizes visual elements and provides real-time multilingual conversion, addressing labor and cost issues in media production, enabling global content distribution.

JP2026073377APending Publication Date: 2026-05-01SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Conventional media production in CG work requires significant labor and costs, and subtitles in video content often lead to missed important information due to limitations, hindering international expansion and effective multilingual support.

Method used

A generative AI model that allows users to upload content data for automatic visual element adjustment and real-time multilingual conversion, optimizing media content for diverse cultural spheres.

Benefits of technology

Enables efficient production and distribution of media content compatible with multiple languages and cultural backgrounds, reducing labor and costs while ensuring all viewers can access critical information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026073377000001_ABST
    Figure 2026073377000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] The means by which users upload content data to the platform, The server analyzes the uploaded content data and identifies visual elements such as color and font, A means by which a generative AI model automatically adjusts the visual elements of content data based on the analysis results, A server automatically converts content data into multiple languages ​​and generates audio and subtitles, A system including means for transmitting adjusted and converted content data to a user's terminal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In conventional CG work in media production, a lot of labor and high costs are required, which has been a factor hindering the international expansion of media. Also, the limitations of subtitles in video content may cause viewers to miss important information. There is a need for means to efficiently optimize visual elements and support multiple languages.

Means for Solving the Problems

[0005] This invention provides a generative AI model that allows users to easily upload content data to a platform, where a server analyzes the data to identify and automatically adjust visual elements such as color and font. Furthermore, it provides a system that supports international expansion by converting content data into multiple languages ​​and generating audio and subtitles in real time. This system is equipped with encoding capabilities and can also optimize based on cultural background and language. This overcomes conventional challenges and enables the efficient production and distribution of media content that is compatible with diverse cultural spheres.

[0006] A "user" is an individual or company that uploads content data to the platform and uses a generative AI system to perform visual and linguistic transformations on media content.

[0007] A "platform" is an online system that provides the infrastructure for users to upload content data and for servers and generative AI models to process it.

[0008] A "server" is a central computing device that stores, analyzes, adjusts, and translates content data received from users into multiple languages.

[0009] "Content data" refers to media data, including video and images, and is digital information that is subject to adjustment of visual elements and multilingual translation.

[0010] "Analysis" is the process by which a server identifies colors, fonts, and other visual elements from content data and extracts the information necessary for processing.

[0011] A "generative AI model" is an algorithm that automatically adjusts the visual elements of content data based on analysis results, optimizing them according to cultural background and language.

[0012] "Multilingual conversion" is the process of translating audio and subtitles within content data into multiple languages ​​specified by the user and generating them in real time.

[0013] "Encoding" is the process of converting content data that has been adjusted and translated into multiple languages ​​into a format that can be distributed.

[0014] "Cultural background" refers to historical, social, and cultural factors that are considered when selecting colors and fonts to suit the audience of a particular region or country. [Brief explanation of the drawing]

[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11]It is a sequence diagram showing the processing flow of the data processing system in Embodiment 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when combined with an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when combined with an emotion engine.

Mode for Carrying Out the Invention

[0016] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, a processor with a reference numeral (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0019] In the following embodiments, a RAM (Random Access Memory) with a reference numeral is a memory in which information is temporarily stored and is used as a work memory by the processor. [[ID=​In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0023] [First Embodiment]

[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0036] This invention provides a system utilizing generative AI to facilitate the international distribution of content data by users. Specific embodiments thereof are described below.

[0037] First, the user accesses the platform from their device, selects the content data they wish to edit, and uploads it. This content data includes media data such as still images and videos, and is sent to the server in the file format specified by the user.

[0038] The server receives the uploaded content data and stores it in secure storage. Here, the server initiates an analysis process, running algorithms to identify the content's color, fonts, and other visual elements. These analysis results are then used as input data for a generative AI model.

[0039] The generative AI model is implemented on the server and automatically optimizes specified visual elements. Specifically, it adjusts the color tone of the video according to the cultural background and changes fonts to those with high legibility. This enables more natural and appropriate content display for viewers from different cultural backgrounds.

[0040] Furthermore, the server performs multilingual translation on content data refined by a generative AI model. It converts audio to text and translates it in real time into multiple specified languages. This translated text is integrated into the video as subtitles and provided to the viewer.

[0041] As a concrete example, consider a case where an educational institution user delivers online lectures internationally. This user uploads lecture videos to the platform and uses a generative AI system to adjust the video brightness and translate the lecture content into multiple languages. This allows students around the world to participate in the classes without experiencing language barriers.

[0042] Finally, the server encodes the adjusted and translated content data and prepares it for transmission to the user's device. Users can download the transmitted content, check its quality, and make corrections as needed. This system streamlines media production and enables cost-effective international distribution.

[0043] The following describes the processing flow.

[0044] Step 1:

[0045] Users access the platform from their devices and select the content data they wish to edit. They then proceed to upload the selected content data to the platform in the specified format.

[0046] Step 2:

[0047] The server receives content data uploaded by users. The received data is stored in secure storage for analysis of its visual elements.

[0048] Step 3:

[0049] The server initiates the process of analyzing the uploaded content data. Using analysis algorithms, it identifies colors, fonts, and other visual elements within the data, and extracts detailed information about these elements.

[0050] Step 4:

[0051] The generative AI model receives the analysis results provided by the server as input data and automatically adjusts them. This optimizes the visual elements of the content data based on specified standards and cultural elements.

[0052] Step 5:

[0053] The server applies a multilingual translation process to content data refined by a generative AI model. It converts audio portions into text information and translates them into multiple languages ​​using real-time speech recognition and translation technologies.

[0054] Step 6:

[0055] The server integrates the translated text into the content data and encodes the adjusted and translated content into a deliverable format. It then prepares this encoded data for transmission to the user.

[0056] Step 7:

[0057] Users download the edited content sent from the server to their devices. They then review the downloaded content and make any necessary final adjustments or corrections to complete the preparation for distribution.

[0058] (Example 1)

[0059] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0060] The problem lies in the lack of effective and appropriate international digital information sharing methods for delivering content to diverse audiences with different cultural backgrounds and languages. This results in visual incongruity and language barriers, making it difficult for content to reach its full potential.

[0061] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0062] In this invention, the server includes means for analyzing digital information to identify visual elements such as color and font, means for a generating AI model to automatically adjust the digital information based on the analysis results, and means for generating audio and subtitles translated into multiple languages. This makes it possible to provide digital information in a form suitable for viewers from different cultural backgrounds and languages.

[0063] A "user" refers to an individual or group that provides or edits digital information within the system.

[0064] "Digital information" refers to media data such as still images and videos that can be sent and received over a network.

[0065] A "network" refers to a communication system that connects computers and devices to each other, enabling the exchange of digital information.

[0066] A "server" refers to a computer system that analyzes, stores, and distributes digital information in response to user requests.

[0067] "Analysis" refers to data processing used to identify visual elements of digital information, such as color tones and fonts.

[0068] "Visual elements" refer to the visual characteristics included in digital information, such as color schemes and fonts.

[0069] A "generative AI model" refers to an interactive computational model that automatically adjusts the visual elements of digital information based on analysis results.

[0070] "Automatic adjustment" refers to the process by which the generated AI model optimizes visual elements based on the analysis results.

[0071] "Multilingual conversion" refers to the process of translating text and audio within digital information into multiple languages.

[0072] "Acoustics" refers to audio data related to digital information.

[0073] "Subtitles" refer to text displayed on the screen as written information that can be read from audio.

[0074] A "display device" refers to the screen of a computer or terminal used to visually output digital information.

[0075] A "prompt statement" refers to an input statement used to instruct a generative AI model to perform a specific process.

[0076] "Visualization" refers to the process of representing the visual characteristics of digital information according to its cultural context.

[0077] "Encoding" refers to the process of converting digital information into a format suitable for distribution and storage.

[0078] As a specific embodiment of this invention, a system is provided in which a user, a terminal, and a server work together to streamline the international sharing of digital information. First, the user accesses an online platform from their terminal. There, the user selects digital information such as still images and videos that they wish to edit and uploads them to the platform in a specified format.

[0079] The server receives the digital information and stores it securely in storage. The server uses image processing software to detect visual elements such as color and font. Specific examples include image processing libraries and text recognition technologies. The data obtained through this analysis is used as input data for automatic adjustment by a generative AI model.

[0080] The generative AI model is implemented on the server and optimizes the visual elements of digital information according to each cultural background. The AI ​​model adjusts the color tone of the digital information and changes fonts to those with high legibility. This ensures that content is delivered in a way that is adaptable to viewers from different cultural backgrounds.

[0081] Furthermore, the server converts the audio to text and translates it into multiple languages ​​in real time using a multilingual translation engine. The translated text is generated as subtitles and integrated into the digital information. As a specific example, natural language processing technology using deep learning is utilized.

[0082] One example of this is when an educational institution user distributes online lectures internationally. The user can upload lecture videos and instruct a generative AI model to adjust the video's color and fonts and translate it into multiple languages. This makes it easier for students around the world to participate in the lectures.

[0083] As an example of a prompt, entering "Adjust the color scheme of the content for the Asian market and translate it into English, Spanish, and French" will provide multicultural digital information.

[0084] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0085] Step 1:

[0086] Users access the platform from their devices, select the digital information they wish to edit, and upload it. The input consists of still images and video data, and the format must be compatible with the platform and server. The output is the specified digital information sent to the server.

[0087] Step 2:

[0088] The server analyzes the received digital information and stores it in secure storage. The input is digital information sent by the user, which is encrypted and stored for enhanced reliability. The output is prepared for the extraction of data for identified visual elements. At this stage, the database is properly indexed to allow for rapid retrieval of information from storage.

[0089] Step 3:

[0090] The server uses an image processing library to analyze the visual elements of digital information in detail, such as color tones and fonts. The digital information saved in step 2 is used as input. Specific operations include pixel analysis using image processing algorithms and text extraction using OCR technology. As output, detailed analysis data of the visual elements becomes input data for the generated AI model.

[0091] Step 4:

[0092] The generative AI model automatically adjusts visual elements based on the analysis data identified within the server. The input is the analysis results from step 3, and the output is optimized color tones and fonts to reflect the cultural context. Specific examples include processes where the generative AI performs hue conversion and font replacement.

[0093] Step 5:

[0094] The server converts audio to text and uses a translation engine to translate the digital information into multiple languages. The input is digital information refined by a generative AI model, and the output is multilingual subtitle text that is integrated into the video data. This translation process uses natural language processing techniques to achieve both accuracy and speed.

[0095] Step 6:

[0096] The server executes the encoding process and prepares the data for transmission to the user's terminal. The input is multilingual translated digital information, which is encoded into a delivery-adaptive format. The output is digital information ready for smooth playback on the user's terminal.

[0097] Step 7:

[0098] The user downloads the digital information sent to their device and checks its quality. The input is encoded digital information received from the server, and the output is a state where the user can view and evaluate the data. Specifically, the user can check the quality of the video and subtitles and request corrections as needed.

[0099] (Application Example 1)

[0100] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0101] When media data created by content creators and users is distributed internationally, cultural backgrounds and language differences present challenges in visual elements and language translation. This makes it difficult to display content naturally and effectively to audiences from different cultures, hindering reach to international audiences.

[0102] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0103] In this invention, the server includes means for a user to upload media data to a communication infrastructure, means for a computer to analyze the uploaded media data and identify visual elements such as color and font, means for a generating AI algorithm to automatically adjust the visual elements of the media data based on the analysis results, means for a computer to convert the automatically adjusted media data into multiple languages ​​and generate audio and subtitles, means for transmitting the adjusted and converted media data to the user's information terminal, and means for converting video footage shot by the user into content optimized according to the cultural background. This enables the provision of natural and appropriate content to viewers in different cultural spheres and facilitates content sharing from an international perspective.

[0104] A "user" refers to an individual or organization that uploads media data to the system.

[0105] "Media data" refers to digital content that includes visual information such as images and videos.

[0106] "Communication infrastructure" refers to the network environment used for uploading media data.

[0107] A "computer" refers to a device that analyzes media data and processes it to identify visual elements.

[0108] A "generative AI algorithm" refers to an artificial intelligence model that automatically adjusts based on analyzed visual elements.

[0109] "Color" refers to the attributes related to color in media data.

[0110] "Typeface" refers to the typeface or font style of text contained in media data.

[0111] "Visual elements" refer to elements such as color and font that affect the appearance of media data.

[0112] "Multilingual conversion" refers to the process of translating audio and text from media data into multiple languages.

[0113] "Subtitles" refer to text data used to document audio content and display it alongside the video.

[0114] "Information terminal" refers to devices such as smartphones and computers that users use to receive media data.

[0115] "Cultural context" refers to the factors that media data takes into account to ensure it is relevant to the recipient's social and cultural background.

[0116] "Optimization" refers to the automated process of adjusting media data to suit different cultures and perspectives.

[0117] The system implementing this invention begins with the user's terminal uploading media data to a server via a communication infrastructure. The server receives the uploaded media data and performs data analysis using a computer. The analysis process includes techniques for identifying visual elements such as color and font using libraries such as OpenCV.

[0118] Next, the generative AI model takes on the role of automatically adjusting the visual elements of the media data based on these analysis results. Here, AI frameworks such as TENSORFLOW® are used to perform optimization processing according to the cultural background. This optimization involves color adjustments, font changes, and other modifications to enable natural and appropriate display for viewers in different cultural spheres and regions.

[0119] Furthermore, to enable multilingual support, the server utilizes the Google® Translate API to convert the audio contained in the media data into text, and then translates it in real time into the specified multiple languages. The generated translated text is then integrated into the video as subtitles.

[0120] As a concrete example, consider a scenario where a travel blogger processes footage of a local festival shot on their smartphone using this system. In this case, the video is automatically adjusted to suit different cultural backgrounds, and subtitles in multiple languages ​​are added, allowing international viewers to enjoy the content without stress.

[0121] The following are some examples of prompts for a generative AI model.

[0122] "Please optimize the color scheme of this video based on its cultural context and translate the text into multiple languages: English, Japanese, and French."

[0123] In this way, users can easily distribute the content they generate to international markets, enabling efficient access to a wide audience.

[0124] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0125] Step 1:

[0126] The user selects media data on their device and uploads it to the server via the communication infrastructure. The input consists of digital files such as images and videos, which are then stored on the server as output. The user operates the upload screen to select the appropriate file format and send the data to the server.

[0127] Step 2:

[0128] The server receives the uploaded media data and begins data analysis using a computer. The input is the media data saved in step 1, and the output is analysis results regarding visual elements such as color and font. This analysis uses image processing techniques such as OpenCV to identify the color and font style.

[0129] Step 3:

[0130] The generative AI model is launched by the server and adjusts the visual elements of the media data based on the analysis results. The input is the analysis results obtained in step 2, and the output is the adjusted media data. This process uses TensorFlow and performs optimization that takes cultural background into account. Specifically, it adjusts the content to be highly visible through color optimization and font changes.

[0131] Step 4:

[0132] The server converts audio to text and performs multilingual translation on the adjusted media data. The input for this step is the adjusted media data, which is the output of step 3, and the output is translated text data. This includes real-time multilingual translation using the Google Translate API and the creation of subtitles in multiple languages.

[0133] Step 5:

[0134] The server encodes the adjusted and translated media data into a deliverable format. The subtitled media data generated in step 4 is used as input, and the encoded media data is obtained as output. The encoding process converts the data into a format that will play smoothly on the user's device.

[0135] Step 6:

[0136] The server ultimately sends the encoded media data to the user's information terminal. The input is the output data from step 5, and the output is a media file playable on the user's terminal. The user can play the received content on their terminal and view it as internationally appropriate content that takes into account different cultural backgrounds.

[0137] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0138] This invention provides a system that dynamically adjusts visual elements and multilingual translations using a generative AI and emotion engine when users deploy content data internationally. Specific embodiments are described below.

[0139] First, users access the platform using their own devices and upload the content data they wish to edit. This data, including videos and images, is sent to the server.

[0140] The server receives the uploaded content data, stores it securely, and then runs an analysis algorithm to identify visual elements such as color and font. Simultaneously, a sentiment engine runs to collect real-time user sentiment information.

[0141] The generative AI model automatically adjusts the visual elements of the content, taking into account analysis results provided by the server and user emotion information obtained by the emotion engine. For example, if the user indicates a relaxed mood, the model will soften the colors and soften the font.

[0142] Furthermore, the server applies a multilingual translation process to the adjusted content data. It converts speech to text in real time and translates it into multiple languages. The emotion engine also operates in this process, optimizing the expression of the translated text according to the user's emotions.

[0143] For example, if a user providing online educational content uses this system, adjustments will be made based on the learner's emotional state. If the learner is focused, the colors and subtitles will become clearer, and adjustments will be made to make the information easier to understand.

[0144] Finally, the server encodes the adjusted and translated content data and prepares it for transmission to the user's device. The user can review the transmitted content and make further adjustments if necessary. By integrating an emotion engine with generative AI, this system can provide a more personalized content experience and lower the barriers to international distribution and use of media content.

[0145] The following describes the processing flow.

[0146] Step 1:

[0147] Users access the platform from their devices, select the content data they wish to edit, and upload it. Here, users verify the file format and prepare the data according to the platform's specifications.

[0148] Step 2:

[0149] The server receives the uploaded content data and stores it in a secure data storage area. After storage, it hands over the data, ready for analysis, to the analysis module.

[0150] Step 3:

[0151] The server begins analyzing the content data. Using an analysis algorithm, it detects the data's color, font, and other visual elements, and prepares the results as input data for a generating AI model.

[0152] Step 4:

[0153] The server's emotion engine receives emotion data from the user's terminal and analyzes it to identify the user's real-time emotional state. This information is used to adjust the visual elements in the adjustment process.

[0154] Step 5:

[0155] The generative AI model combines analysis results provided by the server with user emotion information to automatically adjust the visual elements of the content data. For example, if the user is feeling stressed, the model will calm the colors and adjust the font to make it easier to read.

[0156] Step 6:

[0157] The server performs multilingual translation based on content data adjusted by a generative AI model. It converts audio data into text and applies multilingual translation in real time. At this time, it adjusts the tone of the translation based on information from the emotion engine.

[0158] Step 7:

[0159] The server encodes the adjusted and translated content data and converts it into a deliverable format. This makes the content data available in a format suitable for the user's viewing environment.

[0160] Step 8:

[0161] Users download the edited content sent to their device and review it. They can check the final quality and make additional manual corrections or adjustments as needed.

[0162] (Example 2)

[0163] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0164] In modern society, the demand for multilingual digital content is increasing. However, when adapting content data to different languages ​​and cultural contexts, visual elements and emotional nuances are often lost. Furthermore, manual adjustments are time-consuming and labor-intensive, thus creating a need for automated systems. This invention aims to enable more personalized visual and linguistic adjustments based on the user's emotional state, thereby facilitating international use.

[0165] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0166] In this invention, the server includes means for analyzing digital data and recognizing visual elements, means for a generating AI model to automatically adjust the visual elements based on the analysis results and emotional information, and means for performing multilingual conversion and speech generation. This makes it possible to adjust digital data to match the user's emotional state and display it in multiple languages.

[0167] A "user" refers to an individual or organization that accesses an information system and provides digital data.

[0168] "Digital data" refers to electronically stored information that includes visual elements such as hue and text format.

[0169] An "information processing system" refers to a device or software platform for receiving and analyzing digital data.

[0170] A "server" refers to a computer that receives, stores, analyzes, and performs various processes on digital data.

[0171] "Analysis" refers to information processing operations used to recognize visual elements within digital data.

[0172] "Hue" refers to the visual characteristics related to color within digital data.

[0173] "Character format" refers to the display style and font of characters contained within digital data.

[0174] A "generative AI model" refers to an algorithm that uses artificial intelligence to adjust digital data based on analysis results and emotional information.

[0175] "Emotional information" refers to data that reflects the user's emotional state.

[0176] "Automatic adjustment" refers to the process by which a generative AI model modifies visual elements based on the user's emotional state.

[0177] "Multilingual conversion" refers to the process of translating digital data into different languages.

[0178] "Speech generation" refers to the process of converting text information from digital data into speech data.

[0179] "User's equipment" refers to electronic devices used by a user to receive and display digital data.

[0180] This invention is a system for efficiently providing users with multilingual digital content. The system has the function of analyzing, adjusting, and translating digital data while considering the user's emotional state. Specifically, it is implemented as follows:

[0181] Users access the information processing system using their own devices and provide digital data they wish to edit or translate. This data may include videos and images and is immediately transmitted to the server.

[0182] The server stores the received digital data and runs an analysis program. The analysis recognizes visual elements such as hue and text format. Furthermore, an emotion engine collects real-time emotional information from the user and passes it to a generative AI model.

[0183] The generative AI model automatically adjusts the color tones and font styles of digital data based on the analysis results of visual elements and emotional information provided by the server. This creates content that matches the user's emotional state. For example, if the user is relaxed, the generative AI model will set the hue to a softer tone and change the font style to a softer font.

[0184] Next, the server translates the digital data, which has been refined by the generative AI model, into multiple languages. Here, speech is converted to text, and then translated into multiple languages. The emotion engine also works during this process, optimizing the translated text to best reflect the user's emotions.

[0185] A concrete example is online educational content. When a user uses this system, the content is automatically adjusted according to the learner's emotions, providing a more effective learning experience. For focused learners, the video's hue and subtitles are adjusted to appear more clearly.

[0186] Regarding the use of this system, one example of inputting prompt text into the generating AI model is: "Optimize multilingual translation by adjusting hue and font style based on the user's emotional state."

[0187] In this way, by combining emotional information with generative AI models, it becomes possible to internationally distribute digital data and provide personalized content tailored to the user.

[0188] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0189] Step 1:

[0190] Users access the information processing system through their own devices and select and provide the digital data (such as videos and images) they wish to edit. The input is the digital data specified by the user. The device uses a transfer protocol to send this data to the server. The output is the digital data uploaded to the server.

[0191] Step 2:

[0192] The server securely stores the received digital data in storage and then runs an analysis program. The input is digital data provided by the user. The analysis uses an algorithm to extract visual elements (hue, character format, etc.). The output is data with the visual elements identified.

[0193] Step 3:

[0194] The server uses an emotion engine to collect real-time emotional information from users. The input is raw emotional data that reflects the user's current state. Specifically, it analyzes facial expressions and voice tone acquired from the camera and microphone. The output is quantified emotional information.

[0195] Step 4:

[0196] The generative AI model adjusts the visual elements of digital data based on the analysis results of the provided visual elements and emotional information. The input is the analysis results from step 2 and the emotional information from step 3. The generative AI model follows prompts and performs processing such as softening the hue and adjusting the font format. The output is the adjusted digital data.

[0197] Step 5:

[0198] The server applies multilingual translation to the adjusted digital data. The input is digital data with adjusted visual elements. The server performs audio-to-text conversion and then translates it into various languages. The output is digital data translated into multiple languages.

[0199] Step 6:

[0200] The server encodes the digital data, which has been adjusted and translated, into a format suitable for distribution. The input is digital data translated into multiple languages. The encoding process involves appropriate format conversion. The output is encoded digital data.

[0201] Step 7:

[0202] The server prepares to send the encoded digital data to the user's terminal. The input is data encoded in a deliverable format. The server transmits this data over the network. The output is the digital data sent to the user's terminal.

[0203] (Application Example 2)

[0204] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0205] Traditional content delivery systems have faced challenges in dynamically adjusting visual and linguistic attributes in response to user emotions, making it difficult to provide a personalized viewing experience. Furthermore, even in multilingual translation, simply converting languages ​​often fails to capture cultural and emotional nuances.

[0206] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0207] In this invention, the server includes means for the user to provide information to an information processing platform, means for a data processing device to analyze the provided information and identify visual attributes, and means for an information generation system to automatically adjust the visual attributes based on the analysis results and sentiment analysis results. This enables the dynamic optimization of the viewing experience based on the user's emotions and allows for multilingual translation that is appropriate to the cultural background and emotions.

[0208] A "user" is an entity that provides information to an information processing platform.

[0209] An "information processing platform" is a place that receives and processes information provided by users.

[0210] A "data processing device" is a mechanism for analyzing provided information and identifying visual attributes and other information.

[0211] "Visual attributes" refer to elements that describe visual characteristics such as color and font.

[0212] An "information generation system" is a device that automatically adjusts the visual attributes of information based on analysis results and sentiment analysis results.

[0213] "Emotional analysis results" refer to information that analyzes the user's emotions and displays the results.

[0214] "Automatic adjustment" refers to the system automatically applying the optimal settings based on the user's emotions and visual attributes.

[0215] "Multilingual information" refers to information expressed in various languages.

[0216] "Multilingual translation" is the process of converting information into multiple languages.

[0217] The system for realizing this invention consists of a user's device, a communication network, and a central data processing unit. The user's device can be a smartphone or a head-mounted display, through which the user provides information, including their emotions, to the information processing platform.

[0218] The server functions as an information processing platform, collecting visual and auditory information sent by the user and analyzing it with a data processing device. The data processing device utilizes hardware devices such as cameras and microphones to analyze the user's emotions from their facial expressions and voice. Commonly used tools like Microsoft® Azure® Face API and Google Cloud Speech-to-Text are useful for emotion analysis, allowing for the identification of the user's emotions in real time.

[0219] The generative AI model automatically adjusts visual elements based on the obtained emotion analysis results and visual attribute information. The generative AI model utilizes OpenAI® generative AI technology to optimize visual elements such as color and font according to the user's emotions.

[0220] Furthermore, the server converts the adjusted information into multilingual information. This includes real-time language translation, enabling translations that are appropriate to cultural backgrounds and emotions. Sentiment analysis is also used in the process of converting the acquired audio information into text and translating it into multiple languages, adjusting the nuances of the translation.

[0221] One specific use case is that when a user is watching a movie, the camera can capture their facial expressions during emotionally charged scenes, allowing the system to adjust the screen's color tones slightly and simplify fonts to provide a more emotionally responsive viewing experience.

[0222] An example of a prompt message would be, "Analyze the emotion from the image data of the user crying, and tell me how to optimize the visual elements accordingly." By providing this input to the generating AI model, appropriate visual adjustments will be performed.

[0223] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0224] Step 1:

[0225] The user begins viewing the content using their own device. The user's device captures the user's facial expressions and voice in real time via the camera and microphone, and transmits this data to the server.

[0226] Step 2:

[0227] The server performs sentiment analysis based on the received data. It uses Microsoft Azure's Face API and Google Cloud Speech-to-Text to identify emotional data from the user's facial expressions and voice. In this process, the input is the user's facial expressions and voice data, and the output is data indicating the user's emotional state. Based on the analysis results, numerical data representing the emotions the user is expressing is obtained.

[0228] Step 3:

[0229] The server sends the sentiment analysis results and visual data of the content to the generative AI model. Based on this data, the generative AI model begins the process of optimizing visual elements such as color and font. The input is the sentiment analysis results and initial visual data, and the output is the optimized visual data. The generative AI model automatically generates a visual representation that best suits the user's current emotions.

[0230] Step 4:

[0231] The server performs multilingual translation using optimized visual data. It translates the audio, converted to text in real time, into multiple languages ​​and adjusts the nuances of the translation based on sentiment analysis results. In this step, the input is audio-text data and sentiment data, and the output is multilingual translated text tailored to the user's emotions. The translated text is displayed in a way that is optimal for the user.

[0232] Step 5:

[0233] The server encodes the adjusted visual and translated data and prepares it for transmission to the user's device. The final input is optimized visual and translated data, and the output is encoded data in a deliverable format. The encoded data is played seamlessly on the user's device, providing an optimal viewing experience.

[0234] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0235] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0236] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0237] [Second Embodiment]

[0238] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0239] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0240] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0241] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0242] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0243] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0244] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0245] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0246] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0247] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0248] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0249] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0250] This invention provides a system utilizing generative AI to facilitate the international distribution of content data by users. Specific embodiments thereof are described below.

[0251] First, the user accesses the platform from their device, selects the content data they wish to edit, and uploads it. This content data includes media data such as still images and videos, and is sent to the server in the file format specified by the user.

[0252] The server receives the uploaded content data and stores it in secure storage. Here, the server initiates an analysis process, running algorithms to identify the content's color, fonts, and other visual elements. These analysis results are then used as input data for a generative AI model.

[0253] The generative AI model is implemented on the server and automatically optimizes specified visual elements. Specifically, it adjusts the color tone of the video according to the cultural background and changes fonts to those with high legibility. This enables more natural and appropriate content display for viewers from different cultural backgrounds.

[0254] Furthermore, the server performs multilingual translation on content data refined by a generative AI model. It converts audio to text and translates it in real time into multiple specified languages. This translated text is integrated into the video as subtitles and provided to the viewer.

[0255] As a concrete example, consider a case where an educational institution user delivers online lectures internationally. This user uploads lecture videos to the platform and uses a generative AI system to adjust the video brightness and translate the lecture content into multiple languages. This allows students around the world to participate in the classes without experiencing language barriers.

[0256] Finally, the server encodes the adjusted and translated content data and prepares it for transmission to the user's device. Users can download the transmitted content, check its quality, and make corrections as needed. This system streamlines media production and enables cost-effective international distribution.

[0257] The following describes the processing flow.

[0258] Step 1:

[0259] Users access the platform from their devices and select the content data they wish to edit. They then proceed to upload the selected content data to the platform in the specified format.

[0260] Step 2:

[0261] The server receives content data uploaded by users. The received data is stored in secure storage for analysis of its visual elements.

[0262] Step 3:

[0263] The server initiates the process of analyzing the uploaded content data. Using analysis algorithms, it identifies colors, fonts, and other visual elements within the data, and extracts detailed information about these elements.

[0264] Step 4:

[0265] The generative AI model receives the analysis results provided by the server as input data and automatically adjusts them. This optimizes the visual elements of the content data based on specified standards and cultural elements.

[0266] Step 5:

[0267] The server applies a multilingual translation process to content data refined by a generative AI model. It converts audio portions into text information and translates them into multiple languages ​​using real-time speech recognition and translation technologies.

[0268] Step 6:

[0269] The server integrates the translated text into the content data and encodes the adjusted and translated content into a deliverable format. It then prepares this encoded data for transmission to the user.

[0270] Step 7:

[0271] Users download the edited content sent from the server to their devices. They then review the downloaded content and make any necessary final adjustments or corrections to complete the preparation for distribution.

[0272] (Example 1)

[0273] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0274] The problem lies in the lack of effective and appropriate international digital information sharing methods for delivering content to diverse audiences with different cultural backgrounds and languages. This results in visual incongruity and language barriers, making it difficult for content to reach its full potential.

[0275] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0276] In this invention, the server includes means for analyzing digital information to identify visual elements such as color and font, means for a generating AI model to automatically adjust the digital information based on the analysis results, and means for generating audio and subtitles translated into multiple languages. This makes it possible to provide digital information in a form suitable for viewers from different cultural backgrounds and languages.

[0277] A "user" refers to an individual or group that provides or edits digital information within the system.

[0278] "Digital information" refers to media data such as still images and videos that can be sent and received over a network.

[0279] A "network" refers to a communication system that connects computers and devices to each other, enabling the exchange of digital information.

[0280] "Server" refers to a computer system that analyzes, stores, and distributes digital information in response to requests from users.

[0281] "Analysis" refers to data processing for identifying visual elements such as the color tone and font of digital information.

[0282] "Visual elements" refer to characteristics related to the appearance contained in digital information, such as color tone and font.

[0283] "Generative AI model" refers to an interactive computational model that automatically adjusts the visual elements of digital information based on the analysis results.

[0284] "Automatic adjustment" refers to the process by which the generative AI model optimizes the visual elements based on the analysis results.

[0285] "Multilingual conversion" refers to the process of translating text and audio within digital information into multiple languages.

[0286] "Acoustics" refers to audio data related to digital information.

[0287] "Subtitle" refers to text that is displayed on the screen as character information that can read audio information.

[0288] "Display device" refers to the screen of a computer or terminal for visually outputting digital information.

[0289] "Prompt sentence" refers to an input sentence for instructing a specific process to the generative AI model.

[0290] "Visualization" refers to the process of expressing the visual characteristics of digital information according to the cultural background.

[0291] "Encoding" refers to the process of converting digital information into a format suitable for distribution and storage.

[0292] As a specific embodiment of this invention, a system is provided in which a user, a terminal, and a server work together to streamline the international sharing of digital information. First, the user accesses an online platform from their terminal. There, the user selects digital information such as still images and videos that they wish to edit and uploads them to the platform in a specified format.

[0293] The server receives the digital information and stores it securely in storage. The server uses image processing software to detect visual elements such as color and font. Specific examples include image processing libraries and text recognition technologies. The data obtained through this analysis is used as input data for automatic adjustment by a generative AI model.

[0294] The generative AI model is implemented on the server and optimizes the visual elements of digital information according to each cultural background. The AI ​​model adjusts the color tone of the digital information and changes fonts to those with high legibility. This ensures that content is delivered in a way that is adaptable to viewers from different cultural backgrounds.

[0295] Furthermore, the server converts the audio to text and translates it into multiple languages ​​in real time using a multilingual translation engine. The translated text is generated as subtitles and integrated into the digital information. As a specific example, natural language processing technology using deep learning is utilized.

[0296] One example of this is when an educational institution user distributes online lectures internationally. The user can upload lecture videos and instruct a generative AI model to adjust the video's color and fonts and translate it into multiple languages. This makes it easier for students around the world to participate in the lectures.

[0297] As an example of a prompt, entering "Adjust the color scheme of the content for the Asian market and translate it into English, Spanish, and French" will provide multicultural digital information.

[0298] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0299] Step 1:

[0300] Users access the platform from their devices, select the digital information they wish to edit, and upload it. The input consists of still images and video data, and the format must be compatible with the platform and server. The output is the specified digital information sent to the server.

[0301] Step 2:

[0302] The server analyzes the received digital information and stores it in secure storage. The input is digital information sent by the user, which is encrypted and stored for enhanced reliability. The output is prepared for the extraction of data for identified visual elements. At this stage, the database is properly indexed to allow for rapid retrieval of information from storage.

[0303] Step 3:

[0304] The server uses an image processing library to analyze the visual elements of digital information in detail, such as color tones and fonts. The digital information saved in step 2 is used as input. Specific operations include pixel analysis using image processing algorithms and text extraction using OCR technology. As output, detailed analysis data of the visual elements becomes input data for the generated AI model.

[0305] Step 4:

[0306] The generative AI model automatically adjusts visual elements based on the analyzed data identified within the server. The input is the analysis result of Step 3, and the output is that the color tone and font are optimized in a form corresponding to the cultural background. As a specific example, there is a process where the generative AI performs hue conversion, font replacement, etc.

[0307] Step 5:

[0308] The server converts the voice into text and performs a process of converting digital information into multiple languages using a translation engine. The input is the digital information adjusted by the generative AI model, and as the output, multilingual subtitle text is generated and integrated into the video data. Natural language processing technology is used in this translation process to achieve both accuracy and speed.

[0309] Step 6:

[0310] The server executes an encoding process and prepares to transmit it to the user's terminal. The input is the digitally information translated into multiple languages, which is encoded into a delivery-adaptive format. The output is digital information in a state that can be smoothly played on the user's terminal.

[0311] Step 7:

[0312] <00009�3>The user downloads the digital information transmitted to the terminal and checks its quality. The input is the encoded digital information received from the server, and the output is a state where the user can view and evaluate the data. As a specific operation, the user can check the quality of the video and subtitles and request corrections if necessary.

[0313] (Application Example 1)

[0314] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0315] When media data created by content creators and users is distributed internationally, cultural backgrounds and language differences present challenges in visual elements and language translation. This makes it difficult to display content naturally and effectively to audiences from different cultures, hindering reach to international audiences.

[0316] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0317] In this invention, the server includes means for a user to upload media data to a communication infrastructure, means for a computer to analyze the uploaded media data and identify visual elements such as color and font, means for a generating AI algorithm to automatically adjust the visual elements of the media data based on the analysis results, means for a computer to convert the automatically adjusted media data into multiple languages ​​and generate audio and subtitles, means for transmitting the adjusted and converted media data to the user's information terminal, and means for converting video footage shot by the user into content optimized according to the cultural background. This enables the provision of natural and appropriate content to viewers in different cultural spheres and facilitates content sharing from an international perspective.

[0318] A "user" refers to an individual or organization that uploads media data to the system.

[0319] "Media data" refers to digital content that includes visual information such as images and videos.

[0320] "Communication infrastructure" refers to the network environment used for uploading media data.

[0321] A "computer" refers to a device that analyzes media data and processes it to identify visual elements.

[0322] A "generative AI algorithm" refers to an artificial intelligence model that automatically adjusts based on analyzed visual elements.

[0323] "Color" refers to the attributes related to color in media data.

[0324] "Typeface" refers to the typeface or font style of text contained in media data.

[0325] "Visual elements" refer to elements such as color and font that affect the appearance of media data.

[0326] "Multilingual conversion" refers to the process of translating audio and text from media data into multiple languages.

[0327] "Subtitles" refer to text data used to document audio content and display it alongside the video.

[0328] "Information terminal" refers to devices such as smartphones and computers that users use to receive media data.

[0329] "Cultural context" refers to the factors that media data takes into account to ensure it is relevant to the recipient's social and cultural background.

[0330] "Optimization" refers to the automated process of adjusting media data to suit different cultures and perspectives.

[0331] The system implementing this invention begins with the user's terminal uploading media data to a server via a communication infrastructure. The server receives the uploaded media data and performs data analysis using a computer. The analysis process includes techniques for identifying visual elements such as color and font using libraries such as OpenCV.

[0332] Next, the generative AI model takes on the role of automatically adjusting the visual elements of the media data based on these analysis results. Here, AI frameworks such as TensorFlow are used to perform optimization processing according to the cultural background. This optimization involves color adjustments, font changes, and other modifications to enable natural and appropriate display for viewers from different cultural backgrounds and regions.

[0333] Furthermore, to enable multilingual support, the server utilizes the Google Translate API to convert the audio contained in the media data into text, and then translates it in real time into the specified multiple languages. The generated translated text is then integrated into the video as subtitles.

[0334] As a concrete example, consider a scenario where a travel blogger processes footage of a local festival shot on their smartphone using this system. In this case, the video is automatically adjusted to suit different cultural backgrounds, and subtitles in multiple languages ​​are added, allowing international viewers to enjoy the content without stress.

[0335] The following are some examples of prompts for a generative AI model.

[0336] "Please optimize the color scheme of this video based on its cultural context and translate the text into multiple languages: English, Japanese, and French."

[0337] In this way, users can easily distribute the content they generate to international markets, enabling efficient access to a wide audience.

[0338] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0339] Step 1:

[0340] The user selects media data on their device and uploads it to the server via the communication infrastructure. The input consists of digital files such as images and videos, which are then stored on the server as output. The user operates the upload screen to select the appropriate file format and send the data to the server.

[0341] Step 2:

[0342] The server receives the uploaded media data and begins data analysis using a computer. The input is the media data saved in step 1, and the output is analysis results regarding visual elements such as color and font. This analysis uses image processing techniques such as OpenCV to identify the color and font style.

[0343] Step 3:

[0344] The generative AI model is launched by the server and adjusts the visual elements of the media data based on the analysis results. The input is the analysis results obtained in step 2, and the output is the adjusted media data. This process uses TensorFlow and performs optimization that takes cultural background into account. Specifically, it adjusts the content to be highly visible through color optimization and font changes.

[0345] Step 4:

[0346] The server converts audio to text and performs multilingual translation on the adjusted media data. The input for this step is the adjusted media data, which is the output of step 3, and the output is translated text data. This includes real-time multilingual translation using the Google Translate API and the creation of subtitles in multiple languages.

[0347] Step 5:

[0348] The server encodes the adjusted and translated media data into a deliverable format. The subtitled media data generated in step 4 is used as input, and the encoded media data is obtained as output. The encoding process converts the data into a format that will play smoothly on the user's device.

[0349] Step 6:

[0350] The server ultimately sends the encoded media data to the user's information terminal. The input is the output data from step 5, and the output is a media file playable on the user's terminal. The user can play the received content on their terminal and view it as internationally appropriate content that takes into account different cultural backgrounds.

[0351] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0352] This invention provides a system that dynamically adjusts visual elements and multilingual translations using a generative AI and emotion engine when users deploy content data internationally. Specific embodiments are described below.

[0353] First, users access the platform using their own devices and upload the content data they wish to edit. This data, including videos and images, is sent to the server.

[0354] The server receives the uploaded content data, stores it securely, and then runs an analysis algorithm to identify visual elements such as color and font. Simultaneously, a sentiment engine runs to collect real-time user sentiment information.

[0355] The generative AI model automatically adjusts the visual elements of the content, taking into account analysis results provided by the server and user emotion information obtained by the emotion engine. For example, if the user indicates a relaxed mood, the model will soften the colors and soften the font.

[0356] Furthermore, the server applies a multilingual translation process to the adjusted content data. It converts speech to text in real time and translates it into multiple languages. The emotion engine also operates in this process, optimizing the expression of the translated text according to the user's emotions.

[0357] For example, if a user providing online educational content uses this system, adjustments will be made based on the learner's emotional state. If the learner is focused, the colors and subtitles will become clearer, and adjustments will be made to make the information easier to understand.

[0358] Finally, the server encodes the adjusted and translated content data and prepares it for transmission to the user's device. The user can review the transmitted content and make further adjustments if necessary. By integrating an emotion engine with generative AI, this system can provide a more personalized content experience and lower the barriers to international distribution and use of media content.

[0359] The following describes the processing flow.

[0360] Step 1:

[0361] Users access the platform from their devices, select the content data they wish to edit, and upload it. Here, users verify the file format and prepare the data according to the platform's specifications.

[0362] Step 2:

[0363] The server receives the uploaded content data and stores it in a secure data storage area. After storage, it hands over the data, ready for analysis, to the analysis module.

[0364] Step 3:

[0365] The server begins analyzing the content data. Using an analysis algorithm, it detects the data's color, font, and other visual elements, and prepares the results as input data for a generating AI model.

[0366] Step 4:

[0367] The server's emotion engine receives emotion data from the user's terminal and analyzes it to identify the user's real-time emotional state. This information is used to adjust the visual elements in the adjustment process.

[0368] Step 5:

[0369] The generative AI model combines analysis results provided by the server with user emotion information to automatically adjust the visual elements of the content data. For example, if the user is feeling stressed, the model will calm the colors and adjust the font to make it easier to read.

[0370] Step 6:

[0371] The server performs multilingual translation based on content data adjusted by a generative AI model. It converts audio data into text and applies multilingual translation in real time. At this time, it adjusts the tone of the translation based on information from the emotion engine.

[0372] Step 7:

[0373] The server encodes the adjusted and translated content data and converts it into a deliverable format. This makes the content data available in a format suitable for the user's viewing environment.

[0374] Step 8:

[0375] Users download the edited content sent to their device and review it. They can check the final quality and make additional manual corrections or adjustments as needed.

[0376] (Example 2)

[0377] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0378] In modern society, the demand for multilingual digital content is increasing. However, when adapting content data to different languages ​​and cultural contexts, visual elements and emotional nuances are often lost. Furthermore, manual adjustments are time-consuming and labor-intensive, thus creating a need for automated systems. This invention aims to enable more personalized visual and linguistic adjustments based on the user's emotional state, thereby facilitating international use.

[0379] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0380] In this invention, the server includes means for analyzing digital data and recognizing visual elements, means for a generating AI model to automatically adjust the visual elements based on the analysis results and emotional information, and means for performing multilingual conversion and speech generation. This makes it possible to adjust digital data to match the user's emotional state and display it in multiple languages.

[0381] A "user" refers to an individual or organization that accesses an information system and provides digital data.

[0382] "Digital data" refers to electronically stored information that includes visual elements such as hue and text format.

[0383] An "information processing system" refers to a device or software platform for receiving and analyzing digital data.

[0384] A "server" refers to a computer that receives, stores, analyzes, and performs various processes on digital data.

[0385] "Analysis" refers to information processing operations used to recognize visual elements within digital data.

[0386] "Hue" refers to the visual characteristics related to color within digital data.

[0387] "Character format" refers to the display style and font of characters contained within digital data.

[0388] A "generative AI model" refers to an algorithm that uses artificial intelligence to adjust digital data based on analysis results and emotional information.

[0389] "Emotional information" refers to data that reflects the user's emotional state.

[0390] "Automatic adjustment" refers to the process by which a generative AI model modifies visual elements based on the user's emotional state.

[0391] "Multilingual conversion" refers to the process of translating digital data into different languages.

[0392] "Speech generation" refers to the process of converting text information from digital data into speech data.

[0393] "User's equipment" refers to electronic devices used by a user to receive and display digital data.

[0394] This invention is a system for efficiently providing users with multilingual digital content. The system has the function of analyzing, adjusting, and translating digital data while considering the user's emotional state. Specifically, it is implemented as follows:

[0395] Users access the information processing system using their own devices and provide digital data they wish to edit or translate. This data may include videos and images and is immediately transmitted to the server.

[0396] The server stores the received digital data and runs an analysis program. The analysis recognizes visual elements such as hue and text format. Furthermore, an emotion engine collects real-time emotional information from the user and passes it to a generative AI model.

[0397] The generative AI model automatically adjusts the color tones and font styles of digital data based on the analysis results of visual elements and emotional information provided by the server. This creates content that matches the user's emotional state. For example, if the user is relaxed, the generative AI model will set the hue to a softer tone and change the font style to a softer font.

[0398] Next, the server translates the digital data, which has been refined by the generative AI model, into multiple languages. Here, speech is converted to text, and then translated into multiple languages. The emotion engine also works during this process, optimizing the translated text to best reflect the user's emotions.

[0399] A concrete example is online educational content. When a user uses this system, the content is automatically adjusted according to the learner's emotions, providing a more effective learning experience. For focused learners, the video's hue and subtitles are adjusted to appear more clearly.

[0400] Regarding the use of this system, one example of inputting prompt text into the generating AI model is: "Optimize multilingual translation by adjusting hue and font style based on the user's emotional state."

[0401] In this way, by combining emotional information with generative AI models, it becomes possible to internationally distribute digital data and provide personalized content tailored to the user.

[0402] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0403] Step 1:

[0404] Users access the information processing system through their own devices and select and provide the digital data (such as videos and images) they wish to edit. The input is the digital data specified by the user. The device uses a transfer protocol to send this data to the server. The output is the digital data uploaded to the server.

[0405] Step 2:

[0406] The server securely stores the received digital data in storage and then runs an analysis program. The input is digital data provided by the user. The analysis uses an algorithm to extract visual elements (hue, character format, etc.). The output is data with the visual elements identified.

[0407] Step 3:

[0408] The server uses an emotion engine to collect real-time emotional information from users. The input is raw emotional data that reflects the user's current state. Specifically, it analyzes facial expressions and voice tone acquired from the camera and microphone. The output is quantified emotional information.

[0409] Step 4:

[0410] The generative AI model adjusts the visual elements of digital data based on the analysis results of the provided visual elements and emotional information. The input is the analysis results from step 2 and the emotional information from step 3. The generative AI model follows prompts and performs processing such as softening the hue and adjusting the font format. The output is the adjusted digital data.

[0411] Step 5:

[0412] The server applies multilingual translation to the adjusted digital data. The input is digital data with adjusted visual elements. The server performs audio-to-text conversion and then translates it into various languages. The output is digital data translated into multiple languages.

[0413] Step 6:

[0414] The server encodes the digital data, which has been adjusted and translated, into a format suitable for distribution. The input is digital data translated into multiple languages. The encoding process involves appropriate format conversion. The output is encoded digital data.

[0415] Step 7:

[0416] The server prepares to send the encoded digital data to the user's terminal. The input is data encoded in a deliverable format. The server transmits this data over the network. The output is the digital data sent to the user's terminal.

[0417] (Application Example 2)

[0418] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0419] Traditional content delivery systems have faced challenges in dynamically adjusting visual and linguistic attributes in response to user emotions, making it difficult to provide a personalized viewing experience. Furthermore, even in multilingual translation, simply converting languages ​​often fails to capture cultural and emotional nuances.

[0420] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0421] In this invention, the server includes means for the user to provide information to an information processing platform, means for a data processing device to analyze the provided information and identify visual attributes, and means for an information generation system to automatically adjust the visual attributes based on the analysis results and sentiment analysis results. This enables the dynamic optimization of the viewing experience based on the user's emotions and allows for multilingual translation that is appropriate to the cultural background and emotions.

[0422] A "user" is an entity that provides information to an information processing platform.

[0423] An "information processing platform" is a place that receives and processes information provided by users.

[0424] A "data processing device" is a mechanism for analyzing provided information and identifying visual attributes and other information.

[0425] "Visual attributes" refer to elements that describe visual characteristics such as color and font.

[0426] An "information generation system" is a device that automatically adjusts the visual attributes of information based on analysis results and sentiment analysis results.

[0427] "Emotional analysis results" refer to information that analyzes the user's emotions and displays the results.

[0428] "Automatic adjustment" refers to the system automatically applying the optimal settings based on the user's emotions and visual attributes.

[0429] "Multilingual information" refers to information expressed in various languages.

[0430] "Multilingual translation" is the process of converting information into multiple languages.

[0431] The system for realizing this invention consists of a user's device, a communication network, and a central data processing unit. The user's device can be a smartphone or a head-mounted display, through which the user provides information, including their emotions, to the information processing platform.

[0432] The server functions as an information processing platform, collecting visual and auditory information sent by the user and analyzing it with a data processing device. The data processing device utilizes hardware devices such as cameras and microphones to analyze the user's emotions from their facial expressions and voice. Commonly used tools like Microsoft Azure's Face API and Google Cloud Speech-to-Text are useful for emotion analysis, allowing for the identification of the user's emotions in real time.

[0433] The generative AI model automatically adjusts visual elements based on the obtained emotion analysis results and visual attribute information. The generative AI model utilizes OpenAI's generative AI technology to optimize visual elements such as color and font according to the user's emotions.

[0434] Furthermore, the server converts the adjusted information into multilingual information. This includes real-time language translation, enabling translations that are appropriate to cultural backgrounds and emotions. Sentiment analysis is also used in the process of converting the acquired audio information into text and translating it into multiple languages, adjusting the nuances of the translation.

[0435] One specific use case is that when a user is watching a movie, the camera can capture their facial expressions during emotionally charged scenes, allowing the system to adjust the screen's color tones slightly and simplify fonts to provide a more emotionally responsive viewing experience.

[0436] An example of a prompt message would be, "Analyze the emotion from the image data of the user crying, and tell me how to optimize the visual elements accordingly." By providing this input to the generating AI model, appropriate visual adjustments will be performed.

[0437] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0438] Step 1:

[0439] The user begins viewing the content using their own device. The user's device captures the user's facial expressions and voice in real time via the camera and microphone, and transmits this data to the server.

[0440] Step 2:

[0441] The server performs sentiment analysis based on the received data. It uses Microsoft Azure's Face API and Google Cloud Speech-to-Text to identify emotional data from the user's facial expressions and voice. In this process, the input is the user's facial expressions and voice data, and the output is data indicating the user's emotional state. Based on the analysis results, numerical data representing the emotions the user is expressing is obtained.

[0442] Step 3:

[0443] The server sends the sentiment analysis results and visual data of the content to the generative AI model. Based on this data, the generative AI model begins the process of optimizing visual elements such as color and font. The input is the sentiment analysis results and initial visual data, and the output is the optimized visual data. The generative AI model automatically generates a visual representation that best suits the user's current emotions.

[0444] Step 4:

[0445] The server performs multilingual translation using optimized visual data. It translates the audio, converted to text in real time, into multiple languages ​​and adjusts the nuances of the translation based on sentiment analysis results. In this step, the input is audio-text data and sentiment data, and the output is multilingual translated text tailored to the user's emotions. The translated text is displayed in a way that is optimal for the user.

[0446] Step 5:

[0447] The server encodes the adjusted visual and translated data and prepares it for transmission to the user's device. The final input is optimized visual and translated data, and the output is encoded data in a deliverable format. The encoded data is played seamlessly on the user's device, providing an optimal viewing experience.

[0448] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0449] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0450] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0451] [Third Embodiment]

[0452] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0453] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0454] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0455] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0456] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0457] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0458] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0459] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0460] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0461] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0462] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0463] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0464] This invention provides a system utilizing generative AI to facilitate the international distribution of content data by users. Specific embodiments thereof are described below.

[0465] First, the user accesses the platform from their device, selects the content data they wish to edit, and uploads it. This content data includes media data such as still images and videos, and is sent to the server in the file format specified by the user.

[0466] The server receives the uploaded content data and stores it in secure storage. Here, the server initiates an analysis process, running algorithms to identify the content's color, fonts, and other visual elements. These analysis results are then used as input data for a generative AI model.

[0467] The generative AI model is implemented on the server and automatically optimizes specified visual elements. Specifically, it adjusts the color tone of the video according to the cultural background and changes fonts to those with high legibility. This enables more natural and appropriate content display for viewers from different cultural backgrounds.

[0468] Furthermore, the server performs multilingual translation on content data refined by a generative AI model. It converts audio to text and translates it in real time into multiple specified languages. This translated text is integrated into the video as subtitles and provided to the viewer.

[0469] As a concrete example, consider a case where an educational institution user delivers online lectures internationally. This user uploads lecture videos to the platform and uses a generative AI system to adjust the video brightness and translate the lecture content into multiple languages. This allows students around the world to participate in the classes without experiencing language barriers.

[0470] Finally, the server encodes the adjusted and translated content data and prepares it for transmission to the user's device. Users can download the transmitted content, check its quality, and make corrections as needed. This system streamlines media production and enables cost-effective international distribution.

[0471] The following describes the processing flow.

[0472] Step 1:

[0473] Users access the platform from their devices and select the content data they wish to edit. They then proceed to upload the selected content data to the platform in the specified format.

[0474] Step 2:

[0475] The server receives content data uploaded by users. The received data is stored in secure storage for analysis of its visual elements.

[0476] Step 3:

[0477] The server initiates the process of analyzing the uploaded content data. Using analysis algorithms, it identifies colors, fonts, and other visual elements within the data, and extracts detailed information about these elements.

[0478] Step 4:

[0479] The generative AI model receives the analysis results provided by the server as input data and automatically adjusts them. This optimizes the visual elements of the content data based on specified standards and cultural elements.

[0480] Step 5:

[0481] The server applies a multilingual translation process to content data refined by a generative AI model. It converts audio portions into text information and translates them into multiple languages ​​using real-time speech recognition and translation technologies.

[0482] Step 6:

[0483] The server integrates the translated text into the content data and encodes the adjusted and translated content into a deliverable format. It then prepares this encoded data for transmission to the user.

[0484] Step 7:

[0485] Users download the edited content sent from the server to their devices. They then review the downloaded content and make any necessary final adjustments or corrections to complete the preparation for distribution.

[0486] (Example 1)

[0487] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0488] The problem lies in the lack of effective and appropriate international digital information sharing methods for delivering content to diverse audiences with different cultural backgrounds and languages. This results in visual incongruity and language barriers, making it difficult for content to reach its full potential.

[0489] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0490] In this invention, the server includes means for analyzing digital information to identify visual elements such as color and font, means for a generating AI model to automatically adjust the digital information based on the analysis results, and means for generating audio and subtitles translated into multiple languages. This makes it possible to provide digital information in a form suitable for viewers from different cultural backgrounds and languages.

[0491] A "user" refers to an individual or group that provides or edits digital information within the system.

[0492] "Digital information" refers to media data such as still images and videos that can be sent and received over a network.

[0493] A "network" refers to a communication system that connects computers and devices to each other, enabling the exchange of digital information.

[0494] A "server" refers to a computer system that analyzes, stores, and distributes digital information in response to user requests.

[0495] "Analysis" refers to data processing used to identify visual elements of digital information, such as color tones and fonts.

[0496] "Visual elements" refer to the visual characteristics included in digital information, such as color schemes and fonts.

[0497] A "generative AI model" refers to an interactive computational model that automatically adjusts the visual elements of digital information based on analysis results.

[0498] "Automatic adjustment" refers to the process by which the generated AI model optimizes visual elements based on the analysis results.

[0499] "Multilingual conversion" refers to the process of translating text and audio within digital information into multiple languages.

[0500] "Acoustics" refers to audio data related to digital information.

[0501] "Subtitles" refer to text displayed on the screen as written information that can be read from audio.

[0502] A "display device" refers to the screen of a computer or terminal used to visually output digital information.

[0503] A "prompt statement" refers to an input statement used to instruct a generative AI model to perform a specific process.

[0504] "Visualization" refers to the process of representing the visual characteristics of digital information according to its cultural context.

[0505] "Encoding" refers to the process of converting digital information into a format suitable for distribution and storage.

[0506] As a specific embodiment of this invention, a system is provided in which a user, a terminal, and a server work together to streamline the international sharing of digital information. First, the user accesses an online platform from their terminal. There, the user selects digital information such as still images and videos that they wish to edit and uploads them to the platform in a specified format.

[0507] The server receives the digital information and stores it securely in storage. The server uses image processing software to detect visual elements such as color and font. Specific examples include image processing libraries and text recognition technologies. The data obtained through this analysis is used as input data for automatic adjustment by a generative AI model.

[0508] The generative AI model is implemented on the server and optimizes the visual elements of digital information according to each cultural background. The AI ​​model adjusts the color tone of the digital information and changes fonts to those with high legibility. This ensures that content is delivered in a way that is adaptable to viewers from different cultural backgrounds.

[0509] Furthermore, the server converts the audio to text and translates it into multiple languages ​​in real time using a multilingual translation engine. The translated text is generated as subtitles and integrated into the digital information. As a specific example, natural language processing technology using deep learning is utilized.

[0510] One example of this is when an educational institution user distributes online lectures internationally. The user can upload lecture videos and instruct a generative AI model to adjust the video's color and fonts and translate it into multiple languages. This makes it easier for students around the world to participate in the lectures.

[0511] As an example of a prompt, entering "Adjust the color scheme of the content for the Asian market and translate it into English, Spanish, and French" will provide multicultural digital information.

[0512] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0513] Step 1:

[0514] Users access the platform from their devices, select the digital information they wish to edit, and upload it. The input consists of still images and video data, and the format must be compatible with the platform and server. The output is the specified digital information sent to the server.

[0515] Step 2:

[0516] The server analyzes the received digital information and stores it in secure storage. The input is digital information sent by the user, which is encrypted and stored for enhanced reliability. The output is prepared for the extraction of data for identified visual elements. At this stage, the database is properly indexed to allow for rapid retrieval of information from storage.

[0517] Step 3:

[0518] The server uses an image processing library to analyze the visual elements of digital information in detail, such as color tones and fonts. The digital information saved in step 2 is used as input. Specific operations include pixel analysis using image processing algorithms and text extraction using OCR technology. As output, detailed analysis data of the visual elements becomes input data for the generated AI model.

[0519] Step 4:

[0520] The generative AI model automatically adjusts visual elements based on the analysis data identified within the server. The input is the analysis results from step 3, and the output is optimized color tones and fonts to reflect the cultural context. Specific examples include processes where the generative AI performs hue conversion and font replacement.

[0521] Step 5:

[0522] The server converts audio to text and uses a translation engine to translate the digital information into multiple languages. The input is digital information refined by a generative AI model, and the output is multilingual subtitle text that is integrated into the video data. This translation process uses natural language processing techniques to achieve both accuracy and speed.

[0523] Step 6:

[0524] The server executes the encoding process and prepares the data for transmission to the user's terminal. The input is multilingual translated digital information, which is encoded into a delivery-adaptive format. The output is digital information ready for smooth playback on the user's terminal.

[0525] Step 7:

[0526] The user downloads the digital information sent to their device and checks its quality. The input is encoded digital information received from the server, and the output is a state where the user can view and evaluate the data. Specifically, the user can check the quality of the video and subtitles and request corrections as needed.

[0527] (Application Example 1)

[0528] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0529] When media data created by content creators and users is distributed internationally, cultural backgrounds and language differences present challenges in visual elements and language translation. This makes it difficult to display content naturally and effectively to audiences from different cultures, hindering reach to international audiences.

[0530] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0531] In this invention, the server includes means for a user to upload media data to a communication infrastructure, means for a computer to analyze the uploaded media data and identify visual elements such as color and font, means for a generating AI algorithm to automatically adjust the visual elements of the media data based on the analysis results, means for a computer to convert the automatically adjusted media data into multiple languages ​​and generate audio and subtitles, means for transmitting the adjusted and converted media data to the user's information terminal, and means for converting video footage shot by the user into content optimized according to the cultural background. This enables the provision of natural and appropriate content to viewers in different cultural spheres and facilitates content sharing from an international perspective.

[0532] A "user" refers to an individual or organization that uploads media data to the system.

[0533] "Media data" refers to digital content that includes visual information such as images and videos.

[0534] "Communication infrastructure" refers to the network environment used for uploading media data.

[0535] A "computer" refers to a device that analyzes media data and processes it to identify visual elements.

[0536] A "generative AI algorithm" refers to an artificial intelligence model that automatically adjusts based on analyzed visual elements.

[0537] "Color" refers to the attributes related to color in media data.

[0538] "Typeface" refers to the typeface or font style of text contained in media data.

[0539] "Visual elements" refer to elements such as color and font that affect the appearance of media data.

[0540] "Multilingual conversion" refers to the process of translating audio and text from media data into multiple languages.

[0541] "Subtitles" refer to text data used to document audio content and display it alongside the video.

[0542] "Information terminal" refers to devices such as smartphones and computers that users use to receive media data.

[0543] "Cultural context" refers to the factors that media data takes into account to ensure it is relevant to the recipient's social and cultural background.

[0544] "Optimization" refers to the automated process of adjusting media data to suit different cultures and perspectives.

[0545] The system implementing this invention begins with the user's terminal uploading media data to a server via a communication infrastructure. The server receives the uploaded media data and performs data analysis using a computer. The analysis process includes techniques for identifying visual elements such as color and font using libraries such as OpenCV.

[0546] Next, the generative AI model takes on the role of automatically adjusting the visual elements of the media data based on these analysis results. Here, AI frameworks such as TensorFlow are used to perform optimization processing according to the cultural background. This optimization involves color adjustments, font changes, and other modifications to enable natural and appropriate display for viewers from different cultural backgrounds and regions.

[0547] Furthermore, to enable multilingual support, the server utilizes the Google Translate API to convert the audio contained in the media data into text, and then translates it in real time into the specified multiple languages. The generated translated text is then integrated into the video as subtitles.

[0548] As a concrete example, consider a scenario where a travel blogger processes footage of a local festival shot on their smartphone using this system. In this case, the video is automatically adjusted to suit different cultural backgrounds, and subtitles in multiple languages ​​are added, allowing international viewers to enjoy the content without stress.

[0549] The following are some examples of prompts for a generative AI model.

[0550] "Please optimize the color scheme of this video based on its cultural context and translate the text into multiple languages: English, Japanese, and French."

[0551] In this way, users can easily distribute the content they generate to international markets, enabling efficient access to a wide audience.

[0552] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0553] Step 1:

[0554] The user selects media data on their device and uploads it to the server via the communication infrastructure. The input consists of digital files such as images and videos, which are then stored on the server as output. The user operates the upload screen to select the appropriate file format and send the data to the server.

[0555] Step 2:

[0556] The server receives the uploaded media data and begins data analysis using a computer. The input is the media data saved in step 1, and the output is analysis results regarding visual elements such as color and font. This analysis uses image processing techniques such as OpenCV to identify the color and font style.

[0557] Step 3:

[0558] The generative AI model is launched by the server and adjusts the visual elements of the media data based on the analysis results. The input is the analysis results obtained in step 2, and the output is the adjusted media data. This process uses TensorFlow and performs optimization that takes cultural background into account. Specifically, it adjusts the content to be highly visible through color optimization and font changes.

[0559] Step 4:

[0560] The server converts audio to text and performs multilingual translation on the adjusted media data. The input for this step is the adjusted media data, which is the output of step 3, and the output is translated text data. This includes real-time multilingual translation using the Google Translate API and the creation of subtitles in multiple languages.

[0561] Step 5:

[0562] The server encodes the adjusted and translated media data into a deliverable format. The subtitled media data generated in step 4 is used as input, and the encoded media data is obtained as output. The encoding process converts the data into a format that will play smoothly on the user's device.

[0563] Step 6:

[0564] The server ultimately sends the encoded media data to the user's information terminal. The input is the output data from step 5, and the output is a media file playable on the user's terminal. The user can play the received content on their terminal and view it as internationally appropriate content that takes into account different cultural backgrounds.

[0565] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0566] This invention provides a system that dynamically adjusts visual elements and multilingual translations using a generative AI and emotion engine when users deploy content data internationally. Specific embodiments are described below.

[0567] First, users access the platform using their own devices and upload the content data they wish to edit. This data, including videos and images, is sent to the server.

[0568] The server receives the uploaded content data, stores it securely, and then runs an analysis algorithm to identify visual elements such as color and font. Simultaneously, a sentiment engine runs to collect real-time user sentiment information.

[0569] The generative AI model automatically adjusts the visual elements of the content, taking into account analysis results provided by the server and user emotion information obtained by the emotion engine. For example, if the user indicates a relaxed mood, the model will soften the colors and soften the font.

[0570] Furthermore, the server applies a multilingual translation process to the adjusted content data. It converts speech to text in real time and translates it into multiple languages. The emotion engine also operates in this process, optimizing the expression of the translated text according to the user's emotions.

[0571] For example, if a user providing online educational content uses this system, adjustments will be made based on the learner's emotional state. If the learner is focused, the colors and subtitles will become clearer, and adjustments will be made to make the information easier to understand.

[0572] Finally, the server encodes the adjusted and translated content data and prepares it for transmission to the user's device. The user can review the transmitted content and make further adjustments if necessary. By integrating an emotion engine with generative AI, this system can provide a more personalized content experience and lower the barriers to international distribution and use of media content.

[0573] The following describes the processing flow.

[0574] Step 1:

[0575] Users access the platform from their devices, select the content data they wish to edit, and upload it. Here, users verify the file format and prepare the data according to the platform's specifications.

[0576] Step 2:

[0577] The server receives the uploaded content data and stores it in a secure data storage area. After storage, it hands over the data, ready for analysis, to the analysis module.

[0578] Step 3:

[0579] The server begins analyzing the content data. Using an analysis algorithm, it detects the data's color, font, and other visual elements, and prepares the results as input data for a generating AI model.

[0580] Step 4:

[0581] The server's emotion engine receives emotion data from the user's terminal and analyzes it to identify the user's real-time emotional state. This information is used to adjust the visual elements in the adjustment process.

[0582] Step 5:

[0583] The generative AI model combines analysis results provided by the server with user emotion information to automatically adjust the visual elements of the content data. For example, if the user is feeling stressed, the model will calm the colors and adjust the font to make it easier to read.

[0584] Step 6:

[0585] The server performs multilingual translation based on content data adjusted by a generative AI model. It converts audio data into text and applies multilingual translation in real time. At this time, it adjusts the tone of the translation based on information from the emotion engine.

[0586] Step 7:

[0587] The server encodes the adjusted and translated content data and converts it into a deliverable format. This makes the content data available in a format suitable for the user's viewing environment.

[0588] Step 8:

[0589] Users download the edited content sent to their device and review it. They can check the final quality and make additional manual corrections or adjustments as needed.

[0590] (Example 2)

[0591] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0592] In modern society, the demand for multilingual digital content is increasing. However, when adapting content data to different languages ​​and cultural contexts, visual elements and emotional nuances are often lost. Furthermore, manual adjustments are time-consuming and labor-intensive, thus creating a need for automated systems. This invention aims to enable more personalized visual and linguistic adjustments based on the user's emotional state, thereby facilitating international use.

[0593] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0594] In this invention, the server includes means for analyzing digital data and recognizing visual elements, means for a generating AI model to automatically adjust the visual elements based on the analysis results and emotional information, and means for performing multilingual conversion and speech generation. This makes it possible to adjust digital data to match the user's emotional state and display it in multiple languages.

[0595] A "user" refers to an individual or organization that accesses an information system and provides digital data.

[0596] "Digital data" refers to electronically stored information that includes visual elements such as hue and text format.

[0597] An "information processing system" refers to a device or software platform for receiving and analyzing digital data.

[0598] A "server" refers to a computer that receives, stores, analyzes, and performs various processes on digital data.

[0599] "Analysis" refers to information processing operations used to recognize visual elements within digital data.

[0600] "Hue" refers to the visual characteristics related to color within digital data.

[0601] "Character format" refers to the display style and font of characters contained within digital data.

[0602] A "generative AI model" refers to an algorithm that uses artificial intelligence to adjust digital data based on analysis results and emotional information.

[0603] "Emotional information" refers to data that reflects the user's emotional state.

[0604] "Automatic adjustment" refers to the process by which a generative AI model modifies visual elements based on the user's emotional state.

[0605] "Multilingual conversion" refers to the process of translating digital data into different languages.

[0606] "Speech generation" refers to the process of converting text information from digital data into speech data.

[0607] "User's equipment" refers to electronic devices used by a user to receive and display digital data.

[0608] This invention is a system for efficiently providing users with multilingual digital content. The system has the function of analyzing, adjusting, and translating digital data while considering the user's emotional state. Specifically, it is implemented as follows:

[0609] Users access the information processing system using their own devices and provide digital data they wish to edit or translate. This data may include videos and images and is immediately transmitted to the server.

[0610] The server stores the received digital data and runs an analysis program. The analysis recognizes visual elements such as hue and text format. Furthermore, an emotion engine collects real-time emotional information from the user and passes it to a generative AI model.

[0611] The generative AI model automatically adjusts the color tones and font styles of digital data based on the analysis results of visual elements and emotional information provided by the server. This creates content that matches the user's emotional state. For example, if the user is relaxed, the generative AI model will set the hue to a softer tone and change the font style to a softer font.

[0612] Next, the server translates the digital data, which has been refined by the generative AI model, into multiple languages. Here, speech is converted to text, and then translated into multiple languages. The emotion engine also works during this process, optimizing the translated text to best reflect the user's emotions.

[0613] A concrete example is online educational content. When a user uses this system, the content is automatically adjusted according to the learner's emotions, providing a more effective learning experience. For focused learners, the video's hue and subtitles are adjusted to appear more clearly.

[0614] Regarding the use of this system, one example of inputting prompt text into the generating AI model is: "Optimize multilingual translation by adjusting hue and font style based on the user's emotional state."

[0615] In this way, by combining emotional information with generative AI models, it becomes possible to internationally distribute digital data and provide personalized content tailored to the user.

[0616] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0617] Step 1:

[0618] Users access the information processing system through their own devices and select and provide the digital data (such as videos and images) they wish to edit. The input is the digital data specified by the user. The device uses a transfer protocol to send this data to the server. The output is the digital data uploaded to the server.

[0619] Step 2:

[0620] The server securely stores the received digital data in storage and then runs an analysis program. The input is digital data provided by the user. The analysis uses an algorithm to extract visual elements (hue, character format, etc.). The output is data with the visual elements identified.

[0621] Step 3:

[0622] The server uses an emotion engine to collect real-time emotional information from users. The input is raw emotional data that reflects the user's current state. Specifically, it analyzes facial expressions and voice tone acquired from the camera and microphone. The output is quantified emotional information.

[0623] Step 4:

[0624] The generative AI model adjusts the visual elements of digital data based on the analysis results of the provided visual elements and emotional information. The input is the analysis results from step 2 and the emotional information from step 3. The generative AI model follows prompts and performs processing such as softening the hue and adjusting the font format. The output is the adjusted digital data.

[0625] Step 5:

[0626] The server applies multilingual translation to the adjusted digital data. The input is digital data with adjusted visual elements. The server performs audio-to-text conversion and then translates it into various languages. The output is digital data translated into multiple languages.

[0627] Step 6:

[0628] The server encodes the digital data, which has been adjusted and translated, into a format suitable for distribution. The input is digital data translated into multiple languages. The encoding process involves appropriate format conversion. The output is encoded digital data.

[0629] Step 7:

[0630] The server prepares to send the encoded digital data to the user's terminal. The input is data encoded in a deliverable format. The server transmits this data over the network. The output is the digital data sent to the user's terminal.

[0631] (Application Example 2)

[0632] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0633] Traditional content delivery systems have faced challenges in dynamically adjusting visual and linguistic attributes in response to user emotions, making it difficult to provide a personalized viewing experience. Furthermore, even in multilingual translation, simply converting languages ​​often fails to capture cultural and emotional nuances.

[0634] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0635] In this invention, the server includes means for the user to provide information to an information processing platform, means for a data processing device to analyze the provided information and identify visual attributes, and means for an information generation system to automatically adjust the visual attributes based on the analysis results and sentiment analysis results. This enables the dynamic optimization of the viewing experience based on the user's emotions and allows for multilingual translation that is appropriate to the cultural background and emotions.

[0636] A "user" is an entity that provides information to an information processing platform.

[0637] An "information processing platform" is a place that receives and processes information provided by users.

[0638] A "data processing device" is a mechanism for analyzing provided information and identifying visual attributes and other information.

[0639] "Visual attributes" refer to elements that describe visual characteristics such as color and font.

[0640] An "information generation system" is a device that automatically adjusts the visual attributes of information based on analysis results and sentiment analysis results.

[0641] "Emotional analysis results" refer to information that analyzes the user's emotions and displays the results.

[0642] "Automatic adjustment" refers to the system automatically applying the optimal settings based on the user's emotions and visual attributes.

[0643] "Multilingual information" refers to information expressed in various languages.

[0644] "Multilingual translation" is the process of converting information into multiple languages.

[0645] The system for realizing this invention consists of a user's device, a communication network, and a central data processing unit. The user's device can be a smartphone or a head-mounted display, through which the user provides information, including their emotions, to the information processing platform.

[0646] The server functions as an information processing platform, collecting visual and auditory information sent by the user and analyzing it with a data processing device. The data processing device utilizes hardware devices such as cameras and microphones to analyze the user's emotions from their facial expressions and voice. Commonly used tools like Microsoft Azure's Face API and Google Cloud Speech-to-Text are useful for emotion analysis, allowing for the identification of the user's emotions in real time.

[0647] The generative AI model automatically adjusts visual elements based on the obtained emotion analysis results and visual attribute information. The generative AI model utilizes OpenAI's generative AI technology to optimize visual elements such as color and font according to the user's emotions.

[0648] Furthermore, the server converts the adjusted information into multilingual information. This includes real-time language translation, enabling translations that are appropriate to cultural backgrounds and emotions. Sentiment analysis is also used in the process of converting the acquired audio information into text and translating it into multiple languages, adjusting the nuances of the translation.

[0649] One specific use case is that when a user is watching a movie, the camera can capture their facial expressions during emotionally charged scenes, allowing the system to adjust the screen's color tones slightly and simplify fonts to provide a more emotionally responsive viewing experience.

[0650] An example of a prompt message would be, "Analyze the emotion from the image data of the user crying, and tell me how to optimize the visual elements accordingly." By providing this input to the generating AI model, appropriate visual adjustments will be performed.

[0651] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0652] Step 1:

[0653] The user begins viewing the content using their own device. The user's device captures the user's facial expressions and voice in real time via the camera and microphone, and transmits this data to the server.

[0654] Step 2:

[0655] The server performs sentiment analysis based on the received data. It uses Microsoft Azure's Face API and Google Cloud Speech-to-Text to identify emotional data from the user's facial expressions and voice. In this process, the input is the user's facial expressions and voice data, and the output is data indicating the user's emotional state. Based on the analysis results, numerical data representing the emotions the user is expressing is obtained.

[0656] Step 3:

[0657] The server sends the sentiment analysis results and visual data of the content to the generative AI model. Based on this data, the generative AI model begins the process of optimizing visual elements such as color and font. The input is the sentiment analysis results and initial visual data, and the output is the optimized visual data. The generative AI model automatically generates a visual representation that best suits the user's current emotions.

[0658] Step 4:

[0659] The server performs multilingual translation using optimized visual data. It translates the audio, converted to text in real time, into multiple languages ​​and adjusts the nuances of the translation based on sentiment analysis results. In this step, the input is audio-text data and sentiment data, and the output is multilingual translated text tailored to the user's emotions. The translated text is displayed in a way that is optimal for the user.

[0660] Step 5:

[0661] The server encodes the adjusted visual and translated data and prepares it for transmission to the user's device. The final input is optimized visual and translated data, and the output is encoded data in a deliverable format. The encoded data is played seamlessly on the user's device, providing an optimal viewing experience.

[0662] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0663] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0664] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0665] [Fourth Embodiment]

[0666] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0667] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0668] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0669] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0670] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0671] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0672] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0673] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0674] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0675] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0676] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0677] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0678] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0679] This invention provides a system utilizing generative AI to facilitate the international distribution of content data by users. Specific embodiments thereof are described below.

[0680] First, the user accesses the platform from their device, selects the content data they wish to edit, and uploads it. This content data includes media data such as still images and videos, and is sent to the server in the file format specified by the user.

[0681] The server receives the uploaded content data and stores it in secure storage. Here, the server initiates an analysis process, running algorithms to identify the content's color, fonts, and other visual elements. These analysis results are then used as input data for a generative AI model.

[0682] The generative AI model is implemented on the server and automatically optimizes specified visual elements. Specifically, it adjusts the color tone of the video according to the cultural background and changes fonts to those with high legibility. This enables more natural and appropriate content display for viewers from different cultural backgrounds.

[0683] Furthermore, the server performs multilingual translation on content data refined by a generative AI model. It converts audio to text and translates it in real time into multiple specified languages. This translated text is integrated into the video as subtitles and provided to the viewer.

[0684] As a concrete example, consider a case where an educational institution user delivers online lectures internationally. This user uploads lecture videos to the platform and uses a generative AI system to adjust the video brightness and translate the lecture content into multiple languages. This allows students around the world to participate in the classes without experiencing language barriers.

[0685] Finally, the server encodes the adjusted and translated content data and prepares it for transmission to the user's device. Users can download the transmitted content, check its quality, and make corrections as needed. This system streamlines media production and enables cost-effective international distribution.

[0686] The following describes the processing flow.

[0687] Step 1:

[0688] Users access the platform from their devices and select the content data they wish to edit. They then proceed to upload the selected content data to the platform in the specified format.

[0689] Step 2:

[0690] The server receives content data uploaded by users. The received data is stored in secure storage for analysis of its visual elements.

[0691] Step 3:

[0692] The server initiates the process of analyzing the uploaded content data. Using analysis algorithms, it identifies colors, fonts, and other visual elements within the data, and extracts detailed information about these elements.

[0693] Step 4:

[0694] The generative AI model receives the analysis results provided by the server as input data and automatically adjusts them. This optimizes the visual elements of the content data based on specified standards and cultural elements.

[0695] Step 5:

[0696] The server applies a multilingual translation process to content data refined by a generative AI model. It converts audio portions into text information and translates them into multiple languages ​​using real-time speech recognition and translation technologies.

[0697] Step 6:

[0698] The server integrates the translated text into the content data and encodes the adjusted and translated content into a deliverable format. It then prepares this encoded data for transmission to the user.

[0699] Step 7:

[0700] Users download the edited content sent from the server to their devices. They then review the downloaded content and make any necessary final adjustments or corrections to complete the preparation for distribution.

[0701] (Example 1)

[0702] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0703] The problem lies in the lack of effective and appropriate international digital information sharing methods for delivering content to diverse audiences with different cultural backgrounds and languages. This results in visual incongruity and language barriers, making it difficult for content to reach its full potential.

[0704] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0705] In this invention, the server includes means for analyzing digital information to identify visual elements such as color and font, means for a generating AI model to automatically adjust the digital information based on the analysis results, and means for generating audio and subtitles translated into multiple languages. This makes it possible to provide digital information in a form suitable for viewers from different cultural backgrounds and languages.

[0706] A "user" refers to an individual or group that provides or edits digital information within the system.

[0707] "Digital information" refers to media data such as still images and videos that can be sent and received over a network.

[0708] A "network" refers to a communication system that connects computers and devices to each other, enabling the exchange of digital information.

[0709] A "server" refers to a computer system that analyzes, stores, and distributes digital information in response to user requests.

[0710] "Analysis" refers to data processing used to identify visual elements of digital information, such as color tones and fonts.

[0711] "Visual elements" refer to the visual characteristics included in digital information, such as color schemes and fonts.

[0712] A "generative AI model" refers to an interactive computational model that automatically adjusts the visual elements of digital information based on analysis results.

[0713] "Automatic adjustment" refers to the process by which the generated AI model optimizes visual elements based on the analysis results.

[0714] "Multilingual conversion" refers to the process of translating text and audio within digital information into multiple languages.

[0715] "Acoustics" refers to audio data related to digital information.

[0716] "Subtitles" refer to text displayed on the screen as written information that can be read from audio.

[0717] A "display device" refers to the screen of a computer or terminal used to visually output digital information.

[0718] A "prompt statement" refers to an input statement used to instruct a generative AI model to perform a specific process.

[0719] "Visualization" refers to the process of representing the visual characteristics of digital information according to its cultural context.

[0720] "Encoding" refers to the process of converting digital information into a format suitable for distribution and storage.

[0721] As a specific embodiment of this invention, a system is provided in which a user, a terminal, and a server work together to streamline the international sharing of digital information. First, the user accesses an online platform from their terminal. There, the user selects digital information such as still images and videos that they wish to edit and uploads them to the platform in a specified format.

[0722] The server receives the digital information and stores it securely in storage. The server uses image processing software to detect visual elements such as color and font. Specific examples include image processing libraries and text recognition technologies. The data obtained through this analysis is used as input data for automatic adjustment by a generative AI model.

[0723] The generative AI model is implemented on the server and optimizes the visual elements of digital information according to each cultural background. The AI ​​model adjusts the color tone of the digital information and changes fonts to those with high legibility. This ensures that content is delivered in a way that is adaptable to viewers from different cultural backgrounds.

[0724] Furthermore, the server converts the audio to text and translates it into multiple languages ​​in real time using a multilingual translation engine. The translated text is generated as subtitles and integrated into the digital information. As a specific example, natural language processing technology using deep learning is utilized.

[0725] One example of this is when an educational institution user distributes online lectures internationally. The user can upload lecture videos and instruct a generative AI model to adjust the video's color and fonts and translate it into multiple languages. This makes it easier for students around the world to participate in the lectures.

[0726] As an example of a prompt, entering "Adjust the color scheme of the content for the Asian market and translate it into English, Spanish, and French" will provide multicultural digital information.

[0727] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0728] Step 1:

[0729] Users access the platform from their devices, select the digital information they wish to edit, and upload it. The input consists of still images and video data, and the format must be compatible with the platform and server. The output is the specified digital information sent to the server.

[0730] Step 2:

[0731] The server analyzes the received digital information and stores it in secure storage. The input is digital information sent by the user, which is encrypted and stored for enhanced reliability. The output is prepared for the extraction of data for identified visual elements. At this stage, the database is properly indexed to allow for rapid retrieval of information from storage.

[0732] Step 3:

[0733] The server uses an image processing library to analyze the visual elements of digital information in detail, such as color tones and fonts. The digital information saved in step 2 is used as input. Specific operations include pixel analysis using image processing algorithms and text extraction using OCR technology. As output, detailed analysis data of the visual elements becomes input data for the generated AI model.

[0734] Step 4:

[0735] The generative AI model automatically adjusts visual elements based on the analysis data identified within the server. The input is the analysis results from step 3, and the output is optimized color tones and fonts to reflect the cultural context. Specific examples include processes where the generative AI performs hue conversion and font replacement.

[0736] Step 5:

[0737] The server converts audio to text and uses a translation engine to translate the digital information into multiple languages. The input is digital information refined by a generative AI model, and the output is multilingual subtitle text that is integrated into the video data. This translation process uses natural language processing techniques to achieve both accuracy and speed.

[0738] Step 6:

[0739] The server executes the encoding process and prepares the data for transmission to the user's terminal. The input is multilingual translated digital information, which is encoded into a delivery-adaptive format. The output is digital information ready for smooth playback on the user's terminal.

[0740] Step 7:

[0741] The user downloads the digital information sent to their device and checks its quality. The input is encoded digital information received from the server, and the output is a state where the user can view and evaluate the data. Specifically, the user can check the quality of the video and subtitles and request corrections as needed.

[0742] (Application Example 1)

[0743] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0744] When media data created by content creators and users is distributed internationally, cultural backgrounds and language differences present challenges in visual elements and language translation. This makes it difficult to display content naturally and effectively to audiences from different cultures, hindering reach to international audiences.

[0745] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0746] In this invention, the server includes means for a user to upload media data to a communication infrastructure, means for a computer to analyze the uploaded media data and identify visual elements such as color and font, means for a generating AI algorithm to automatically adjust the visual elements of the media data based on the analysis results, means for a computer to convert the automatically adjusted media data into multiple languages ​​and generate audio and subtitles, means for transmitting the adjusted and converted media data to the user's information terminal, and means for converting video footage shot by the user into content optimized according to the cultural background. This enables the provision of natural and appropriate content to viewers in different cultural spheres and facilitates content sharing from an international perspective.

[0747] A "user" refers to an individual or organization that uploads media data to the system.

[0748] "Media data" refers to digital content that includes visual information such as images and videos.

[0749] "Communication infrastructure" refers to the network environment used for uploading media data.

[0750] A "computer" refers to a device that analyzes media data and processes it to identify visual elements.

[0751] A "generative AI algorithm" refers to an artificial intelligence model that automatically adjusts based on analyzed visual elements.

[0752] "Color" refers to the attributes related to color in media data.

[0753] "Typeface" refers to the typeface or font style of text contained in media data.

[0754] "Visual elements" refer to elements such as color and font that affect the appearance of media data.

[0755] "Multilingual conversion" refers to the process of translating audio and text from media data into multiple languages.

[0756] "Subtitles" refer to text data used to document audio content and display it alongside the video.

[0757] "Information terminal" refers to devices such as smartphones and computers that users use to receive media data.

[0758] "Cultural context" refers to the factors that media data takes into account to ensure it is relevant to the recipient's social and cultural background.

[0759] "Optimization" refers to the automated process of adjusting media data to suit different cultures and perspectives.

[0760] The system implementing this invention begins with the user's terminal uploading media data to a server via a communication infrastructure. The server receives the uploaded media data and performs data analysis using a computer. The analysis process includes techniques for identifying visual elements such as color and font using libraries such as OpenCV.

[0761] Next, the generative AI model takes on the role of automatically adjusting the visual elements of the media data based on these analysis results. Here, AI frameworks such as TensorFlow are used to perform optimization processing according to the cultural background. This optimization involves color adjustments, font changes, and other modifications to enable natural and appropriate display for viewers from different cultural backgrounds and regions.

[0762] Furthermore, to enable multilingual support, the server utilizes the Google Translate API to convert the audio contained in the media data into text, and then translates it in real time into the specified multiple languages. The generated translated text is then integrated into the video as subtitles.

[0763] As a concrete example, consider a scenario where a travel blogger processes footage of a local festival shot on their smartphone using this system. In this case, the video is automatically adjusted to suit different cultural backgrounds, and subtitles in multiple languages ​​are added, allowing international viewers to enjoy the content without stress.

[0764] The following are some examples of prompts for a generative AI model.

[0765] "Please optimize the color scheme of this video based on its cultural context and translate the text into multiple languages: English, Japanese, and French."

[0766] In this way, users can easily distribute the content they generate to international markets, enabling efficient access to a wide audience.

[0767] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0768] Step 1:

[0769] The user selects media data on their device and uploads it to the server via the communication infrastructure. The input consists of digital files such as images and videos, which are then stored on the server as output. The user operates the upload screen to select the appropriate file format and send the data to the server.

[0770] Step 2:

[0771] The server receives the uploaded media data and begins data analysis using a computer. The input is the media data saved in step 1, and the output is analysis results regarding visual elements such as color and font. This analysis uses image processing techniques such as OpenCV to identify the color and font style.

[0772] Step 3:

[0773] The generative AI model is launched by the server and adjusts the visual elements of the media data based on the analysis results. The input is the analysis results obtained in step 2, and the output is the adjusted media data. This process uses TensorFlow and performs optimization that takes cultural background into account. Specifically, it adjusts the content to be highly visible through color optimization and font changes.

[0774] Step 4:

[0775] The server converts audio to text and performs multilingual translation on the adjusted media data. The input for this step is the adjusted media data, which is the output of step 3, and the output is translated text data. This includes real-time multilingual translation using the Google Translate API and the creation of subtitles in multiple languages.

[0776] Step 5:

[0777] The server encodes the adjusted and translated media data into a deliverable format. The subtitled media data generated in step 4 is used as input, and the encoded media data is obtained as output. The encoding process converts the data into a format that will play smoothly on the user's device.

[0778] Step 6:

[0779] The server ultimately sends the encoded media data to the user's information terminal. The input is the output data from step 5, and the output is a media file playable on the user's terminal. The user can play the received content on their terminal and view it as internationally appropriate content that takes into account different cultural backgrounds.

[0780] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0781] This invention provides a system that dynamically adjusts visual elements and multilingual translations using a generative AI and emotion engine when users deploy content data internationally. Specific embodiments are described below.

[0782] First, users access the platform using their own devices and upload the content data they wish to edit. This data, including videos and images, is sent to the server.

[0783] The server receives the uploaded content data, stores it securely, and then runs an analysis algorithm to identify visual elements such as color and font. Simultaneously, a sentiment engine runs to collect real-time user sentiment information.

[0784] The generative AI model automatically adjusts the visual elements of the content, taking into account analysis results provided by the server and user emotion information obtained by the emotion engine. For example, if the user indicates a relaxed mood, the model will soften the colors and soften the font.

[0785] Furthermore, the server applies a multilingual translation process to the adjusted content data. It converts speech to text in real time and translates it into multiple languages. The emotion engine also operates in this process, optimizing the expression of the translated text according to the user's emotions.

[0786] For example, if a user providing online educational content uses this system, adjustments will be made based on the learner's emotional state. If the learner is focused, the colors and subtitles will become clearer, and adjustments will be made to make the information easier to understand.

[0787] Finally, the server encodes the adjusted and translated content data and prepares it for transmission to the user's device. The user can review the transmitted content and make further adjustments if necessary. By integrating an emotion engine with generative AI, this system can provide a more personalized content experience and lower the barriers to international distribution and use of media content.

[0788] The following describes the processing flow.

[0789] Step 1:

[0790] Users access the platform from their devices, select the content data they wish to edit, and upload it. Here, users verify the file format and prepare the data according to the platform's specifications.

[0791] Step 2:

[0792] The server receives the uploaded content data and stores it in a secure data storage area. After storage, it hands over the data, ready for analysis, to the analysis module.

[0793] Step 3:

[0794] The server begins analyzing the content data. Using an analysis algorithm, it detects the data's color, font, and other visual elements, and prepares the results as input data for a generating AI model.

[0795] Step 4:

[0796] The server's emotion engine receives emotion data from the user's terminal and analyzes it to identify the user's real-time emotional state. This information is used to adjust the visual elements in the adjustment process.

[0797] Step 5:

[0798] The generative AI model combines analysis results provided by the server with user emotion information to automatically adjust the visual elements of the content data. For example, if the user is feeling stressed, the model will calm the colors and adjust the font to make it easier to read.

[0799] Step 6:

[0800] The server performs multilingual translation based on content data adjusted by a generative AI model. It converts audio data into text and applies multilingual translation in real time. At this time, it adjusts the tone of the translation based on information from the emotion engine.

[0801] Step 7:

[0802] The server encodes the adjusted and translated content data and converts it into a deliverable format. This makes the content data available in a format suitable for the user's viewing environment.

[0803] Step 8:

[0804] Users download the edited content sent to their device and review it. They can check the final quality and make additional manual corrections or adjustments as needed.

[0805] (Example 2)

[0806] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0807] In modern society, the demand for multilingual digital content is increasing. However, when adapting content data to different languages ​​and cultural contexts, visual elements and emotional nuances are often lost. Furthermore, manual adjustments are time-consuming and labor-intensive, thus creating a need for automated systems. This invention aims to enable more personalized visual and linguistic adjustments based on the user's emotional state, thereby facilitating international use.

[0808] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0809] In this invention, the server includes means for analyzing digital data and recognizing visual elements, means for a generating AI model to automatically adjust the visual elements based on the analysis results and emotional information, and means for performing multilingual conversion and speech generation. This makes it possible to adjust digital data to match the user's emotional state and display it in multiple languages.

[0810] A "user" refers to an individual or organization that accesses an information system and provides digital data.

[0811] "Digital data" refers to electronically stored information that includes visual elements such as hue and text format.

[0812] An "information processing system" refers to a device or software platform for receiving and analyzing digital data.

[0813] A "server" refers to a computer that receives, stores, analyzes, and performs various processes on digital data.

[0814] "Analysis" refers to information processing operations used to recognize visual elements within digital data.

[0815] "Hue" refers to the visual characteristics related to color within digital data.

[0816] "Character format" refers to the display style and font of characters contained within digital data.

[0817] A "generative AI model" refers to an algorithm that uses artificial intelligence to adjust digital data based on analysis results and emotional information.

[0818] "Emotional information" refers to data that reflects the user's emotional state.

[0819] "Automatic adjustment" refers to the process by which a generative AI model modifies visual elements based on the user's emotional state.

[0820] "Multilingual conversion" refers to the process of translating digital data into different languages.

[0821] "Speech generation" refers to the process of converting text information from digital data into speech data.

[0822] "User's equipment" refers to electronic devices used by a user to receive and display digital data.

[0823] This invention is a system for efficiently providing users with multilingual digital content. The system has the function of analyzing, adjusting, and translating digital data while considering the user's emotional state. Specifically, it is implemented as follows:

[0824] Users access the information processing system using their own devices and provide digital data they wish to edit or translate. This data may include videos and images and is immediately transmitted to the server.

[0825] The server stores the received digital data and runs an analysis program. The analysis recognizes visual elements such as hue and text format. Furthermore, an emotion engine collects real-time emotional information from the user and passes it to a generative AI model.

[0826] The generative AI model automatically adjusts the color tones and font styles of digital data based on the analysis results of visual elements and emotional information provided by the server. This creates content that matches the user's emotional state. For example, if the user is relaxed, the generative AI model will set the hue to a softer tone and change the font style to a softer font.

[0827] Next, the server translates the digital data, which has been refined by the generative AI model, into multiple languages. Here, speech is converted to text, and then translated into multiple languages. The emotion engine also works during this process, optimizing the translated text to best reflect the user's emotions.

[0828] A concrete example is online educational content. When a user uses this system, the content is automatically adjusted according to the learner's emotions, providing a more effective learning experience. For focused learners, the video's hue and subtitles are adjusted to appear more clearly.

[0829] Regarding the use of this system, one example of inputting prompt text into the generating AI model is: "Optimize multilingual translation by adjusting hue and font style based on the user's emotional state."

[0830] In this way, by combining emotional information with generative AI models, it becomes possible to internationally distribute digital data and provide personalized content tailored to the user.

[0831] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0832] Step 1:

[0833] Users access the information processing system through their own devices and select and provide the digital data (such as videos and images) they wish to edit. The input is the digital data specified by the user. The device uses a transfer protocol to send this data to the server. The output is the digital data uploaded to the server.

[0834] Step 2:

[0835] The server securely stores the received digital data in storage and then runs an analysis program. The input is digital data provided by the user. The analysis uses an algorithm to extract visual elements (hue, character format, etc.). The output is data with the visual elements identified.

[0836] Step 3:

[0837] The server uses an emotion engine to collect real-time emotional information from users. The input is raw emotional data that reflects the user's current state. Specifically, it analyzes facial expressions and voice tone acquired from the camera and microphone. The output is quantified emotional information.

[0838] Step 4:

[0839] The generative AI model adjusts the visual elements of digital data based on the analysis results of the provided visual elements and emotional information. The input is the analysis results from step 2 and the emotional information from step 3. The generative AI model follows prompts and performs processing such as softening the hue and adjusting the font format. The output is the adjusted digital data.

[0840] Step 5:

[0841] The server applies multilingual translation to the adjusted digital data. The input is digital data with adjusted visual elements. The server performs audio-to-text conversion and then translates it into various languages. The output is digital data translated into multiple languages.

[0842] Step 6:

[0843] The server encodes the digital data, which has been adjusted and translated, into a format suitable for distribution. The input is digital data translated into multiple languages. The encoding process involves appropriate format conversion. The output is encoded digital data.

[0844] Step 7:

[0845] The server prepares to send the encoded digital data to the user's terminal. The input is data encoded in a deliverable format. The server transmits this data over the network. The output is the digital data sent to the user's terminal.

[0846] (Application Example 2)

[0847] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0848] Traditional content delivery systems have faced challenges in dynamically adjusting visual and linguistic attributes in response to user emotions, making it difficult to provide a personalized viewing experience. Furthermore, even in multilingual translation, simply converting languages ​​often fails to capture cultural and emotional nuances.

[0849] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0850] In this invention, the server includes means for the user to provide information to an information processing platform, means for a data processing device to analyze the provided information and identify visual attributes, and means for an information generation system to automatically adjust the visual attributes based on the analysis results and sentiment analysis results. This enables the dynamic optimization of the viewing experience based on the user's emotions and allows for multilingual translation that is appropriate to the cultural background and emotions.

[0851] A "user" is an entity that provides information to an information processing platform.

[0852] An "information processing platform" is a place that receives and processes information provided by users.

[0853] A "data processing device" is a mechanism for analyzing provided information and identifying visual attributes and other information.

[0854] "Visual attributes" refer to elements that describe visual characteristics such as color and font.

[0855] An "information generation system" is a device that automatically adjusts the visual attributes of information based on analysis results and sentiment analysis results.

[0856] "Emotional analysis results" refer to information that analyzes the user's emotions and displays the results.

[0857] "Automatic adjustment" refers to the system automatically applying the optimal settings based on the user's emotions and visual attributes.

[0858] "Multilingual information" refers to information expressed in various languages.

[0859] "Multilingual translation" is the process of converting information into multiple languages.

[0860] The system for realizing this invention consists of a user's device, a communication network, and a central data processing unit. The user's device can be a smartphone or a head-mounted display, through which the user provides information, including their emotions, to the information processing platform.

[0861] The server functions as an information processing platform, collecting visual and auditory information sent by the user and analyzing it with a data processing device. The data processing device utilizes hardware devices such as cameras and microphones to analyze the user's emotions from their facial expressions and voice. Commonly used tools like Microsoft Azure's Face API and Google Cloud Speech-to-Text are useful for emotion analysis, allowing for the identification of the user's emotions in real time.

[0862] The generative AI model automatically adjusts visual elements based on the obtained emotion analysis results and visual attribute information. The generative AI model utilizes OpenAI's generative AI technology to optimize visual elements such as color and font according to the user's emotions.

[0863] Furthermore, the server converts the adjusted information into multilingual information. This includes real-time language translation, enabling translations that are appropriate to cultural backgrounds and emotions. Sentiment analysis is also used in the process of converting the acquired audio information into text and translating it into multiple languages, adjusting the nuances of the translation.

[0864] One specific use case is that when a user is watching a movie, the camera can capture their facial expressions during emotionally charged scenes, allowing the system to adjust the screen's color tones slightly and simplify fonts to provide a more emotionally responsive viewing experience.

[0865] An example of a prompt message would be, "Analyze the emotion from the image data of the user crying, and tell me how to optimize the visual elements accordingly." By providing this input to the generating AI model, appropriate visual adjustments will be performed.

[0866] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0867] Step 1:

[0868] The user begins viewing the content using their own device. The user's device captures the user's facial expressions and voice in real time via the camera and microphone, and transmits this data to the server.

[0869] Step 2:

[0870] The server performs sentiment analysis based on the received data. It uses Microsoft Azure's Face API and Google Cloud Speech-to-Text to identify emotional data from the user's facial expressions and voice. In this process, the input is the user's facial expressions and voice data, and the output is data indicating the user's emotional state. Based on the analysis results, numerical data representing the emotions the user is expressing is obtained.

[0871] Step 3:

[0872] The server sends the sentiment analysis results and visual data of the content to the generative AI model. Based on this data, the generative AI model begins the process of optimizing visual elements such as color and font. The input is the sentiment analysis results and initial visual data, and the output is the optimized visual data. The generative AI model automatically generates a visual representation that best suits the user's current emotions.

[0873] Step 4:

[0874] The server performs multilingual translation using optimized visual data. It translates the audio, converted to text in real time, into multiple languages ​​and adjusts the nuances of the translation based on sentiment analysis results. In this step, the input is audio-text data and sentiment data, and the output is multilingual translated text tailored to the user's emotions. The translated text is displayed in a way that is optimal for the user.

[0875] Step 5:

[0876] The server encodes the adjusted visual and translated data and prepares it for transmission to the user's device. The final input is optimized visual and translated data, and the output is encoded data in a deliverable format. The encoded data is played seamlessly on the user's device, providing an optimal viewing experience.

[0877] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0878] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0879] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0880] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0881] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0882] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0883] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0884] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0885] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0886] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0887] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0888] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0889] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0890] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0891] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0892] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0893] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0894] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0895] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0896] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0897] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0898] The following is further disclosed regarding the embodiments described above.

[0899] (Claim 1)

[0900] The means by which users upload content data to the platform,

[0901] The server analyzes the uploaded content data and identifies visual elements such as color and font,

[0902] A means by which a generative AI model automatically adjusts the visual elements of content data based on the analysis results,

[0903] A server automatically converts content data into multiple languages ​​and generates audio and subtitles,

[0904] A system including means for transmitting adjusted and converted content data to a user's terminal.

[0905] (Claim 2)

[0906] The system according to claim 1, further comprising means for the server to encode the adjusted and converted content data into a deliverable format.

[0907] (Claim 3)

[0908] The system according to claim 1, comprising means for optimizing the generative AI model based on cultural background and language.

[0909] "Example 1"

[0910] (Claim 1)

[0911] A means for users to upload digital information to a network,

[0912] A server analyzes uploaded digital information and identifies visual elements such as color and font style,

[0913] A means by which a generative AI model automatically adjusts the visual elements of digital information based on the analysis results,

[0914] A server automatically converts digitally adjusted information into multiple languages ​​and generates audio and subtitles,

[0915] Means for transmitting adjusted and converted digital information to the user's display device,

[0916] A means by which the server performs visualization according to the cultural background based on the specified prompt message,

[0917] A system including means for encoding adjusted digital information into a format adapted to the device.

[0918] (Claim 2)

[0919] The system according to claim 1, wherein the server provides processing for encoding the adjusted and converted digital information into a format that can be distributed.

[0920] (Claim 3)

[0921] The system according to claim 1, which includes processing for optimizing the generated AI model to suit cultural background and language.

[0922] "Application Example 1"

[0923] (Claim 1)

[0924] A means for users to upload media data to the communication infrastructure,

[0925] A means of a computer analyzing uploaded media data to identify visual elements such as color and font,

[0926] A means by which a generative AI algorithm automatically adjusts the visual elements of media data based on the analysis results,

[0927] A means for a computer to automatically convert media data into multiple languages ​​and generate audio and subtitles,

[0928] A means for transmitting the adjusted and converted media data to the user's information terminal,

[0929] A system that includes means for converting user-recorded video footage into content optimized according to its cultural context.

[0930] (Claim 2)

[0931] The system according to claim 1, further comprising means for converting the adjusted and converted media data into a format that can be distributed.

[0932] (Claim 3)

[0933] The system according to claim 1, further comprising means for automatically generating multiple subtitles by optimizing the generation AI algorithm based on cultural background and language.

[0934] "Example 2 of combining an emotion engine"

[0935] (Claim 1)

[0936] Means by which users provide digital data to information processing systems,

[0937] The server analyzes the provided digital data and has means to recognize the visual elements of hue and character format,

[0938] A means by which a generative AI model automatically adjusts the visual elements of digital data based on analysis results and emotional information,

[0939] A server automatically converts digitally adjusted data into multiple languages ​​and generates audio and text.

[0940] A system including means for transmitting adjusted and converted digital data to a user's device.

[0941] (Claim 2)

[0942] The system according to claim 1, comprising means for a server to encode the adjusted and converted digital data into a format that can be distributed.

[0943] (Claim 3)

[0944] The system according to claim 1, wherein the generative AI model is equipped with means for performing optimization based on cultural background and language.

[0945] "Application example 2 when combining with an emotional engine"

[0946] (Claim 1)

[0947] Means by which users provide information to an information processing platform,

[0948] A data processing device analyzes the provided information and has means for identifying visual attributes,

[0949] The information generation system provides means for automatically adjusting visual attributes based on analysis results and emotion analysis results,

[0950] A data processing device converts the adjusted information into multilingual information and generates audio and text information,

[0951] Means for transmitting the adjusted and converted information to the user's device,

[0952] A means of analyzing user emotions and dynamically optimizing the viewing experience accordingly,

[0953] A system that includes this.

[0954] (Claim 2)

[0955] The system according to claim 1, wherein the data processing device comprises means for converting the adjusted and converted information into a reproducible format.

[0956] (Claim 3)

[0957] The system according to claim 1, comprising means for performing optimization based on cultural background and language, and for performing multilingual translation appropriate to emotions, for the information generation system. [Explanation of Symbols]

[0958] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. The means by which users upload content data to the platform, The server analyzes the uploaded content data and identifies visual elements such as color and font, A means by which a generative AI model automatically adjusts the visual elements of content data based on the analysis results, A server automatically converts content data into multiple languages ​​and generates audio and subtitles, A system including means for transmitting adjusted and converted content data to a user's terminal.

2. The system according to claim 1, further comprising means for the server to encode the adjusted and converted content data into a format that can be distributed.

3. The system according to claim 1, comprising means for optimizing the generative AI model based on cultural background and language.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A