system

The system addresses event management challenges by using a user interface, natural language processing, multilingual translation, and a virtual presenter to provide flexible and engaging event moderation.

JP2026073403APending Publication Date: 2026-05-01SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Event organizers face challenges in managing events efficiently due to language barriers, cultural differences, and sudden schedule changes, requiring significant time and resources, and existing systems are inadequate in providing flexible and cost-effective solutions.

Method used

A system that includes a user interface for intuitive event information input, a natural language processing model to generate a progress script, multilingual translation, and a virtual presenter with integrated speech synthesis and visual representation, allowing for culturally appropriate and flexible event management.

Benefits of technology

Enables efficient, multilingual, and cost-effective event moderation, accommodating diverse cultural backgrounds and real-time adjustments, providing a realistic and engaging experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026073403000001_ABST
    Figure 2026073403000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means of providing an intuitive user interface for users to input basic event information, A means for generating an event progress script using a natural language processing model based on received information, A multilingual translation method for translating the generated script into multiple languages, A means of presenting the translated script to the user and providing a user-editable interface, A means for controlling a virtual presenter by integrating speech synthesis and visual representation based on an edited script, A means of reflecting real-time information updates during the event, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Event organizers and organizations need to spend a lot of time and effort on prior consultations with the host, and it is difficult to cope with language barriers and sudden schedule changes, and cost reduction is required. Furthermore, flexible hosting according to diverse cultural backgrounds and specific needs is required, but there is a problem that the current system cannot respond to these requirements quickly and efficiently.

Means for Solving the Problems

[0005] This invention solves the above problems by providing a user interface that allows for intuitive input of basic event information, and by receiving that information and automatically generating a progress script using a natural language processing model. Furthermore, it translates the generated script into multiple languages, allows users to easily edit and customize it, and controls a virtual presenter that integrates speech synthesis and visual representation based on the edited script in real time. This enables diverse expressions that are culturally appropriate and meet the individual needs of users, realizing rapid and efficient event management.

[0006] A "user interface" is a screen display and input method that allows a user to input information into a system, perform operations, and receive feedback.

[0007] A "natural language processing model" is a set of algorithms and techniques used to enable computers to understand, analyze, and generate human language.

[0008] A "proceeding script" refers to the script or structure used for hosting an event, including the order in which actions should be taken and the content of what should be said.

[0009] A "multilingual translation method" refers to a technology and system for translating information input in one language into multiple other languages.

[0010] A "virtual presenter" is a computer-generated virtual character with visual and auditory features who acts as the host or moderator for an event.

[0011] "Speech synthesis" is a technology that generates artificial speech based on input text.

[0012] "Visual expression" refers to the means of conveying information visually and the content of that expression, and in particular, to expressions that include animation and video. [Brief explanation of the drawing]

[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0014] An example of an embodiment of the system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0017] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0018] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0019] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0021] [First Embodiment]

[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0034] This invention provides a virtual presenter system for efficient and flexible event hosting. The system includes a user interface, a natural language processing model, multilingual translation capabilities, a virtual presenter, and integrated speech synthesis and visual representation functions. The specific operation flow is described below.

[0035] The user acts as the event organizer, using the terminal's user interface to input information related to the event. As a concrete example, consider information about an international conference. In this case, the user would input the event title, date, participant list, and a brief overview of the proceedings.

[0036] The entered information is sent to the server. The server receives this information and uses a natural language processing model to generate an event progress script. The generated script includes the flow of the moderator's speech and what to say, enabling smooth guidance during the event.

[0037] The server translates the progress script into multiple languages ​​as needed, enabling it to be accessible to international audiences. The translated script is then adjusted for culturally natural expression.

[0038] The translated and adjusted script is presented to the user on their device. The user can review this script and edit or customize specific parts as needed. For example, they can edit the details when introducing a particular speaker.

[0039] Finally, a virtual presenter guides the event using speech synthesis and visuals based on an edited script. The voice and gestures are synchronized at each stage, providing the audience with a realistic, presenter-like experience. The system can also accommodate real-time schedule changes during the event.

[0040] This configuration enables efficient, multilingual, and cost-effective moderation of events. This implementation is widely applicable to various types of events, and is particularly effective in international and multicultural settings.

[0041] The following describes the processing flow.

[0042] Step 1:

[0043] Users enter the necessary information for the event using the user interface on their device. This information includes the event name, date and time, participant information, and schedule.

[0044] Step 2:

[0045] The terminal sends information entered by the user to the server. The transmitted data is received by the server for processing.

[0046] Step 3:

[0047] The server analyzes the received information and automatically generates an event progress script using a natural language processing model. The script is created based on the input information and includes utterances and comments.

[0048] Step 4:

[0049] The server-generated scripts are fed into a translation system for multilingual support. The scripts are translated into the required languages ​​and culturally appropriate, with adjustments made to each language version.

[0050] Step 5:

[0051] The device presents the translated script to the user, who then reviews the content. The user edits and customizes the script for any parts that need correction.

[0052] Step 6:

[0053] The server receives the user's edited script and begins integrating the speech synthesis system and visual representation. It then sets up the virtual presenter's presentation, combining voice, facial expressions, and gestures.

[0054] Step 7:

[0055] On the day of the event, the device will begin the event proceedings with virtual presenters. It will output audio and animations in real time according to the script, and immediately reflect updates to information as needed.

[0056] This series of steps allows the system to manage events efficiently and flexibly.

[0057] (Example 1)

[0058] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0059] Efficiently and multilingually managing international and multicultural events is challenging. Furthermore, events with frequent schedule changes require rapid updates to the program. Existing technologies often require significant human resources and time to address these issues, resulting in high costs.

[0060] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0061] In this invention, the server includes means for providing an intuitive input device for users to input information, means for generating a progress statement using a language model based on the received information, and means for providing a multilingual conversion device for converting the generated statement into multiple languages. This enables smooth hosting in multiple languages ​​and allows for flexible response to information updates during the event.

[0062] A "user" is someone who operates the system, inputs information related to an event, and takes on the role of facilitating its progress.

[0063] An "input device" refers to a device or interface that a user uses to input information.

[0064] A "language model" is a natural language processing technique used to generate a continuous sentence based on input information.

[0065] A "proceeding document" refers to the script or record used within an event, and contains information necessary for the smooth running of the event.

[0066] A "translation device" refers to a technology or means used to translate a generated text into multiple different languages.

[0067] "Multilingual support" means the ability to support multiple languages.

[0068] A "virtual character" is a digital character that integrates sound generation and visual expression to conduct events on screen.

[0069] "Information updates" refer to the immediate reflection of any changes or new information that arise as the event progresses.

[0070] This invention realizes a virtual MC system for efficiently managing events, in which a server, terminal, and user work together.

[0071] First, the user uses the device's intuitive input device to enter information related to the event. Specifically, they can enter the event title, date, participant information, and a summary of the agenda. An example of a prompt message would be: "Please create a progress script for the next event. The title is 'International Conference on Technological Innovation,' the date is November 15, 2023, and the participants are experts from around the world."

[0072] The terminal sends this input information to the server. The server uses a language model based on the received information to generate a continuous sentence. This language model is based on natural language processing technology, and it is conceivable to use existing technologies such as Google's TENSORFLOW® or OpenAI's GPT®.

[0073] The server translates the generated proceedings into multiple languages. Machine translation systems are used for each language, with platforms such as Microsoft® Translator and Amazon Translate available. This multilingual support ensures smooth proceedings for audiences at international events.

[0074] The translated script is sent to the user's device, where they can review its content and make adjustments or edits as needed. For example, it's easy to add or modify names and job titles when introducing specific speakers.

[0075] Ultimately, the server sends the user-edited script to a virtual character, integrating sound generation and visual presentation to run the event. The server uses Google's Text-to-Speech or Amazon Polly for speech synthesis software, and real-time 3D platforms such as Unity or Unreal Engine for visual effects. This allows the virtual character to behave like a real presenter, providing an engaging experience for the audience.

[0076] As described above, this invention combines various components to achieve efficient and flexible event management.

[0077] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0078] Step 1:

[0079] Users enter basic event information using the device's intuitive input interface. This includes the event title, date, participant list, agenda, and language selection. The entered information is stored on the device as digital data.

[0080] Step 2:

[0081] The terminal sends event information received from the user to the server. The information is transferred using a secure communication protocol (e.g., HTTPS), and the server receives this data. The server analyzes the received data and verifies its accuracy.

[0082] Step 3:

[0083] The server generates a sequence of events using a generative AI model based on the received event information. The language model performs natural language processing on the given prompt sentence and creates a script that progresses within the appropriate context. In this process, context is extracted from the input data (event information), and sentences that follow the flow are output.

[0084] Step 4:

[0085] The server translates the generated progress text into multiple specified languages ​​using a multilingual translation device. An automatic translation system is used for each language. The output is a translated script containing culturally appropriate expressions for each language.

[0086] Step 5:

[0087] The server sends the translated script to the terminal. The terminal displays this script to the user, who can review and edit its contents as needed. Specific actions include modifying, adding, or deleting specific phrases. The edited script is saved to the terminal.

[0088] Step 6:

[0089] The server sends the final edited script to the virtual character. The virtual character uses specialized sound generation software to synthesize speech and a system that displays visual representations in real time. In this process, the virtual character conducts the event in a natural manner, like a real presenter, based on the script edited by the user.

[0090] (Application Example 1)

[0091] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0092] Modern events increasingly require support for diverse cultural backgrounds and multilingualism, while simultaneously demanding the provision of realistic experiences in virtual environments. However, efficiently conveying information and managing participant interaction is extremely difficult. Similarly, in product introductions at virtual stores, effectively communicating the value of new products and maintaining participant interest is crucial. Traditional systems have faced challenges in adequately conveying information due to language barriers and cultural differences.

[0093] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0094] In this invention, the server includes means for providing an intuitive user interface for users to input basic event information, means for generating an event progress script using a natural language processing model based on the received information, and multilingual translation means for translating the generated script into multiple languages. This makes it possible to convey information to participants in a multilingual and culturally appropriate manner even in a virtual environment, and to provide participants with an immersive experience through product introductions in a virtual store.

[0095] A "user interface" is an intuitive and easy-to-use screen for users to input information into a system.

[0096] A "natural language processing model" is an algorithm that analyzes input text data to understand and generate human language.

[0097] An "event progress script" is a planned script of presentations designed to ensure an event runs smoothly.

[0098] A "multilingual translation method" is a function that accurately translates information across multiple different languages.

[0099] "Speech synthesis" is a technology that converts text information into speech that sounds like a human voice.

[0100] "Visual expression" refers to methods of conveying information visually through images and animations.

[0101] A "virtual presenter" is a software agent that acts as a digitally constructed moderator, providing information and event management support to the audience.

[0102] "Smart glasses" are wearable devices that can display virtual information in real time.

[0103] A "virtual store" is a digital store environment that is accessible to users online.

[0104] "Real-time information updates" is a feature that allows new information to be reflected instantly as the event progresses.

[0105] To realize this application, the system operates by integrating multiple technical components. The server is primarily responsible for data processing and generation. First, the user inputs basic event information into the terminal through an intuitive user interface. This interface is designed to be easy to operate and to quickly provide the user with the information they need.

[0106] The server uses the received information to automatically generate an event progress script, utilizing a natural language processing model. This generated script is presented to participants in real time within the virtual store using wearable devices such as smart glasses. The generated script is translated into multiple languages ​​using a multilingual translation system. This enables appropriate and effective information transmission to international participants with diverse cultural backgrounds.

[0107] This script is presented by a virtual presenter that integrates speech synthesis technology and visual representation. This presenter provides a natural experience for the audience, updating information in real time as the event progresses. For example, as a user wearing smart glasses walks around a virtual store, it can provide introductory videos of new products along with audio. It can also instantly generate relevant information in text using a generative AI model in response to user actions and questions.

[0108] A concrete example is a scenario where a visual explaining the mechanism of a new product appears on the display of smart glasses, and the user requests more information using voice commands, at which point a virtual presenter explains the product's features and applications in detail. An example of a generated AI prompt is, "Please write a script that explains the product's appeal from an angle that the user is likely to find interesting." In this way, the system can provide participants with a deep sense of immersion and useful information.

[0109] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0110] Step 1:

[0111] Users enter basic event information using their devices. This information includes the event title, date, participant list, and agenda. This information can be entered intuitively through the user interface, and the device then sends it to the server.

[0112] Step 2:

[0113] The server uses natural language processing models to generate an event progress script based on the received event information. The input here is the event information, and the output is the initial draft of the script. This process involves analyzing the elements of the information and compiling the appropriate progress and important points into the script.

[0114] Step 3:

[0115] The generated script is translated into multiple languages ​​on the server by a multilingual translation engine. The input is the script in progress, and the output is the translated versions of the script. The translated script is adjusted to preserve the cultural nuances appropriate for each language.

[0116] Step 4:

[0117] Users can view the translated script on their device and edit its content using the interface. The input is the translated script, and the output is the final script edited by the user. The edited script is then sent back to the server.

[0118] Step 5:

[0119] The server uses the final script to integrate speech synthesis and visual representation, and controls the actions of the virtual presenter. The input is the final script, and the output is the virtual moderator's control commands for the presentation. This provides real-time narration and visual effects in accordance with the progress of the event.

[0120] Step 6:

[0121] Users wearing smart glasses are guided through a virtual store. Specific product information and explanations are presented in response to a script and real-time voice commands from the user. Input is user interaction, and output is product information presented as visual and audio data.

[0122] Step 7:

[0123] During the event, if any new information or changes arise, the server updates the script in real time and immediately reflects them in the virtual presenter. The input is real-time event information, and the output is the updated progress script. This step allows for flexible response to unexpected changes.

[0124] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0125] This invention provides a more advanced interactive experience by combining a virtual presenter system with an emotion engine that recognizes user emotions. The system includes a user interface, a natural language processing model, multilingual translation capabilities, a virtual presenter, integrated speech synthesis and visual representation functions, and an emotion engine. The specific operation flow is described below.

[0126] The user inputs event-related information through their device. This includes basic information as well as conditional settings depending on the situation. This information is sent to the server, which automatically generates an event progress script based on the information.

[0127] The generated script is translated into the required language through a multilingual translation system and adjusted to be a natural expression with cultural relevance. The translated script is then reviewed by the user on their device and edited and customized as needed.

[0128] A key feature here is the use of an emotion engine when the server prepares the virtual presenter's speech synthesis and visual presentation based on the edited script. The emotion engine analyzes the emotions of users and audience members in real time during the event. For example, if a user is nervous, the emotion engine recognizes this and adjusts the system to proceed in a more relaxed tone.

[0129] The analysis results from the emotion engine are transmitted to the server, which then dynamically adjusts the content of the script and the virtual presenter's expressions during the event. For example, if there are signs that the audience's interest is waning, the program will change the presenter's tone and gestures to try and make the presentation more engaging.

[0130] Ultimately, the terminal controls the event proceedings by a virtual presenter, using speech synthesis and visual representations to provide the audience with a realistic hosting experience. The system dynamically updates based on scripts and sentiment analysis results, allowing it to flexibly respond to unexpected situations during the event.

[0131] In this configuration, the system of the present invention can provide an interactive and personalized event experience that meets the needs of users and audiences.

[0132] The following describes the processing flow.

[0133] Step 1:

[0134] The user enters event details (e.g., event name, date, participant information, program details) using the user interface on their device. The device then formats the entered information appropriately and sends it to the server.

[0135] Step 2:

[0136] The server analyzes the information received from the user and generates an event progress script using a natural language processing model. The script includes introductions for each speaker and a description of the event flow in natural language.

[0137] Step 3:

[0138] The server translates the generated script into the required language via a multilingual translation system. The translated script is then adjusted to ensure culturally natural expression.

[0139] Step 4:

[0140] The device presents the translated script to the user. The user can review the script through the interface and edit or modify it as needed.

[0141] Step 5:

[0142] The server receives the edited script and prepares it for speech synthesis and visual representation. During this process, it activates the emotion engine and analyzes the emotions of the user or audience in real time.

[0143] Step 6:

[0144] The emotion engine analyzes the emotions of users and audience members and communicates the results to the server. Based on the information received, the server adjusts the script during the event and changes the tone and gestures of the virtual presenter.

[0145] Step 7:

[0146] During the event, the device controls the virtual presenter, using synthesized speech and animated visuals to manage the event. It also incorporates dynamic expressions to maintain audience engagement, taking into account analysis results from an emotion engine.

[0147] This series of processes allows the system to provide an interactive and personalized event experience, and to flexibly adapt to any situation.

[0148] (Example 2)

[0149] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0150] Modern events are becoming increasingly diverse, requiring personalized experiences that meet the needs of participants. However, traditional methods have struggled to adapt to participants' emotions and cultural backgrounds in real time, making it difficult to create engaging and enjoyable events.

[0151] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0152] In this invention, the server includes means for providing an intuitive interface using a terminal for users to input information, means for generating a progress script using a data processing model based on the received information, and means for performing sentiment analysis and dynamically adjusting the facilitator's expressions according to the state of the audience. This makes it possible to provide a personalized interactive event experience that is tailored to the emotions and cultural backgrounds of the participants.

[0153] A "terminal" is an electronic device that allows a user to input information and interact with a system through an interface.

[0154] An "interface" is a screen or means available to a user for entering information into a system or for reviewing and editing scripts.

[0155] A "data processing model" is an algorithm or system that automatically generates event progress scripts based on received information.

[0156] A "procedure script" is a set of instructions that details the steps and content necessary to ensure an event runs smoothly.

[0157] "Multilingual translation means" refers to a technology or system used to automatically convert a generated progress script into multiple languages.

[0158] A "virtual presenter" is a computer-generated character that fulfills the role of a host or navigator, represented in a digital format.

[0159] "Speech generation and visual representation" refers to technologies that generate speech content as sound for a virtual presenter and provide visual effects for the audience.

[0160] "Emotion analysis" is the process of evaluating the emotional state of users or audiences in real time using sensors and algorithms.

[0161] This system includes terminals for user information input, a server that generates and controls progress scripts, and a mechanism for sentiment analysis. Users first input basic event information and participant characteristics via the terminal. The terminal provides an intuitive interface, making it easy for users to input information.

[0162] Information sent from the terminal is received by the server. The server uses a generative AI model to generate a natural language script based on the received information. The generated script is then translated into the required language using a multilingual translation tool. Tone and cultural appropriateness are also adjusted during this process.

[0163] For example, a possible prompt to input into a generative AI model is: "Generate a script for a 30-minute technical seminar. The target audience is young people, and it should include interactive elements."

[0164] The translated script is returned to the terminal, where the user can review and edit it as needed. Based on the user's edited script, the server adjusts the virtual presenter's voice generation and visual presentation, and prepares for the event. During this process, the server uses an emotion analysis engine to analyze the emotions of the user and audience in real time as the event progresses. Based on the emotion analysis results, the virtual presenter's expressions and tone are dynamically changed.

[0165] In this way, the system can provide users and audiences with a customized, interactive event experience. The system can also flexibly respond to unexpected situations during the event and make dynamic adjustments to ensure a realistic and engaging presentation.

[0166] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0167] Step 1:

[0168] Users use their devices to enter detailed information about the event. This information includes the date and time of the event, the theme, and the profile of the target audience. The entered information is organized as digital data on the device and sent to the server.

[0169] Step 2:

[0170] The server begins data processing based on the information received from the terminal. Specifically, it uses a generative AI model to generate a script for event progression. This model analyzes the input data based on the provided prompts and outputs natural language text. The generated script is then formatted as digital text data.

[0171] Step 3:

[0172] The server sends the generated script to a multilingual translation system. This system translates the script into the target language, evaluates and adjusts for cultural appropriateness. The translated script is then sent back to the terminal. During translation, language data conversion processing is performed, and the output is a natural expression in the target language.

[0173] Step 4:

[0174] The user reviews and edits the translated script on their device. If the user changes terminology in the script or adds new data, the device formats the edited data for resending to the server. This step involves text editing of the data.

[0175] Step 5:

[0176] The server prepares the virtual presenter based on the edited script. Specifically, it uses a speech generation engine to convert the script into audio data and a visual representation system to generate a visual presentation. In parallel, it uses an emotion analysis engine to analyze the emotions of users and the audience and generates data for real-time adjustments.

[0177] Step 6:

[0178] The terminal controls the event host in real time, receiving updated data from the server and dynamically adjusting the presentation. The system combines audio and visuals to create a realistic and interactive experience for the audience. Processing is performed to change the host's tone and gestures in response to the event situation and user feedback.

[0179] (Application Example 2)

[0180] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0181] In modern consumer activity, users require information about products and services that is tailored to their emotions and interests at that moment. However, conventional systems often lack sufficient real-time sentiment analysis and dynamic adjustment of information delivery, limiting the user experience. Furthermore, efficiently and culturally appropriate information delivery to global consumers who require multilingual support remains a challenge.

[0182] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0183] In this invention, the server includes means for providing an intuitive data input interface for users to input basic event information; means for generating an event progress script using a language processing device based on the received information; a translation device for translating the generated script into multiple languages; means for presenting the translated script to the user and providing a user-editable interface; means for controlling a virtual presenter by integrating speech synthesis and visual representation based on the edited script; an emotion analysis device for analyzing customer emotions and dynamically adjusting the information provided; and means for reflecting real-time information updates during the event. This enables personalized information provision in response to user emotions, thereby improving the user experience.

[0184] A "user" is someone who seeks information about products and services during their consumption activities.

[0185] A "data entry interface" is a means for users to intuitively input event and product information.

[0186] A "language processing device" is a device that analyzes and processes natural language and generates event progress scripts in various languages ​​based on the underlying information.

[0187] A "translation device" is a device used to translate a generated script into multiple languages.

[0188] An "interface" is a means by which the user can view and edit the provided translation script.

[0189] "Speech synthesis" is a technology that generates speech based on a processed script.

[0190] "Visual representation" refers to a means of visually representing a virtual narrator in conjunction with speech synthesis.

[0191] A "virtual facilitator" is a virtual display character that plays the role of guiding an event in a digital environment.

[0192] An "emotion analysis device" is a device that analyzes a user's emotions and dynamically adjusts the way information is provided based on that data.

[0193] "Personalization" refers to providing customized information tailored to the individual user's interests and emotions.

[0194] "Real-time" refers to something that happens immediately in the present moment.

[0195] To implement this invention, a server plays a central role. This server operates in conjunction with various software, as described below. First, the user inputs basic event information from a terminal through a data input interface. The terminal used here could be a general-purpose computer or smart device. Next, the server uses a language processing unit to perform natural language processing based on the input information and generates an event progress script. In this process, the server utilizes an advanced generative AI model.

[0196] The generated scripts are translated into multiple languages ​​using a translation tool. Services such as the Google Cloud Translation API can be used for this translation. The translated scripts are presented to the user and can be edited via a dedicated interface. This editing process also takes place on the user's device.

[0197] Once editing is complete, the server combines speech synthesis and visual representation technologies to generate a virtual presenter. Software such as Amazon Polly is used for speech synthesis, and the virtual presenter provides information to the user or audience. Furthermore, the server uses an emotion analyzer to analyze the emotions of the user or audience in real time. This analysis is then processed using machine learning libraries such as TensorFlow.

[0198] Based on the analysis results, the server dynamically adjusts how information is delivered to provide personalized information. For example, if the sentiment analyzer determines that the user seems uninterested while receiving a product description, the system will provide additional interesting information.

[0199] An example of a prompt message could be input to the generating AI model, such as, "Provide more interesting information based on the user's emotions." In this way, the present invention utilizes real-time user feedback to achieve effective information transmission.

[0200] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0201] Step 1:

[0202] The user enters basic event information into a data input interface via their device. The entered information is sent to the server as text data. The server receives this information and prepares it for natural language processing.

[0203] Step 2:

[0204] The server generates an event progress script using a language processing unit based on the received information. In this process, it utilizes a generative AI model to perform advanced natural language processing. The input is basic information provided by the user, and the output is the text data of the integrated event progress script.

[0205] Step 3:

[0206] The server uses a translation tool to translate the generated script into the required languages. The Google Cloud Translation API can be used for this purpose. The input is the text data of the generated script, and the output is the multilingual text data of the translated script.

[0207] Step 4:

[0208] The translated script is displayed on the user's terminal and becomes editable. The user reviews the information through the presented interface and makes corrections as needed. The input is the translated script, and the output is the script edited by the user.

[0209] Step 5:

[0210] Based on the edited script, the server integrates speech synthesis and visual representation to generate a virtual narrator. Amazon Polly or similar tools are used for speech synthesis, and the generated virtual narrator is displayed on the terminal. The input is the user-edited script, and the output is the virtual narrator's visual and audio data.

[0211] Step 6:

[0212] The server uses an emotion analysis device to analyze the emotions of users and audiences in real time. This analysis uses data acquired from the camera and microphone on the device, and is performed using TensorFlow or similar software. The input is the acquired emotion data, and the output is the emotion analysis result.

[0213] Step 7:

[0214] Based on the analysis results, the server dynamically adjusts how information is delivered. Specifically, the information and tone are changed according to the user's emotions. For example, the prompt "Please provide additional information to maintain the user's interest" is input into the generating AI model, and the customized information output is presented by a virtual presenter.

[0215] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0216] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0217] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0218] [Second Embodiment]

[0219] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0220] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0221] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0222] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0223] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0224] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0225] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0226] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0227] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0228] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0229] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0230] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0231] This invention provides a virtual presenter system for efficient and flexible event hosting. The system includes a user interface, a natural language processing model, multilingual translation capabilities, a virtual presenter, and integrated speech synthesis and visual representation functions. The specific operation flow is described below.

[0232] The user acts as the event organizer, using the terminal's user interface to input information related to the event. As a concrete example, consider information about an international conference. In this case, the user would input the event title, date, participant list, and a brief overview of the proceedings.

[0233] The entered information is sent to the server. The server receives this information and uses a natural language processing model to generate an event progress script. The generated script includes the flow of the moderator's speech and what to say, enabling smooth guidance during the event.

[0234] The server translates the progress script into multiple languages ​​as needed, enabling it to be accessible to international audiences. The translated script is then adjusted for culturally natural expression.

[0235] The translated and adjusted script is presented to the user on their device. The user can review this script and edit or customize specific parts as needed. For example, they can edit the details when introducing a particular speaker.

[0236] Finally, a virtual presenter guides the event using speech synthesis and visuals based on an edited script. The voice and gestures are synchronized at each stage, providing the audience with a realistic, presenter-like experience. The system can also accommodate real-time schedule changes during the event.

[0237] This configuration enables efficient, multilingual, and cost-effective moderation of events. This implementation is widely applicable to various types of events, and is particularly effective in international and multicultural settings.

[0238] The following describes the processing flow.

[0239] Step 1:

[0240] Users enter the necessary information for the event using the user interface on their device. This information includes the event name, date and time, participant information, and schedule.

[0241] Step 2:

[0242] The terminal sends information entered by the user to the server. The transmitted data is received by the server for processing.

[0243] Step 3:

[0244] The server analyzes the information it receives and automatically generates an event progress script using a natural language processing model. The script is created based on the input information and includes utterances and comments.

[0245] Step 4:

[0246] The server-generated scripts are fed into a translation system for multilingual support. The scripts are translated into the required languages ​​and culturally appropriate, with adjustments made to each language version.

[0247] Step 5:

[0248] The device presents the translated script to the user, who then reviews the content. The user edits and customizes the script for any parts that need correction.

[0249] Step 6:

[0250] The server receives the user's edited script and begins integrating the speech synthesis system and visual representation. It then sets up the virtual presenter's presentation, combining voice, facial expressions, and gestures.

[0251] Step 7:

[0252] On the day of the event, the device will begin the event proceedings with virtual presenters. It will output audio and animations in real time according to the script, and immediately reflect updates to information as needed.

[0253] This series of steps allows the system to manage events efficiently and flexibly.

[0254] (Example 1)

[0255] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0256] Efficiently and multilingually managing international and multicultural events is challenging. Furthermore, events with frequent schedule changes require rapid updates to the program. Existing technologies often require significant human resources and time to address these issues, resulting in high costs.

[0257] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0258] In this invention, the server includes means for providing an intuitive input device for users to input information, means for generating a progress statement using a language model based on the received information, and means for providing a multilingual conversion device for converting the generated statement into multiple languages. This enables smooth hosting in multiple languages ​​and allows for flexible response to information updates during the event.

[0259] A "user" is someone who operates the system, inputs information related to an event, and takes on the role of facilitating its progress.

[0260] An "input device" refers to a device or interface that a user uses to input information.

[0261] A "language model" is a natural language processing technique used to generate a continuous sentence based on input information.

[0262] A "proceeding document" refers to the script or record used within an event, and contains information necessary for the smooth running of the event.

[0263] A "translation device" refers to a technology or means used to translate a generated text into multiple different languages.

[0264] "Multilingual support" means the ability to support multiple languages.

[0265] A "virtual character" is a digital character that integrates sound generation and visual expression to conduct events on screen.

[0266] "Information updates" refer to the immediate reflection of any changes or new information that arise as the event progresses.

[0267] This invention realizes a virtual MC system for efficiently managing events, in which a server, terminal, and user work together.

[0268] First, the user uses the device's intuitive input device to enter information related to the event. Specifically, they can enter the event title, date, participant information, and a summary of the agenda. An example of a prompt message would be: "Please create a progress script for the next event. The title is 'International Conference on Technological Innovation,' the date is November 15, 2023, and the participants are experts from around the world."

[0269] The terminal sends this input information to the server. The server uses a language model based on the received information to generate a continuous sentence. This language model is based on natural language processing technology, and it is conceivable that existing technologies such as Google's TensorFlow or OpenAI's GPT could be used.

[0270] The server translates the generated proceedings into multiple languages. Machine translation systems are used for each language, with platforms such as Microsoft Translator and Amazon Translate available. This multilingual support ensures smooth proceedings for audiences at international events.

[0271] The translated script is sent to the user's device, where they can review its content and make adjustments or edits as needed. For example, it's easy to add or modify names and job titles when introducing specific speakers.

[0272] Ultimately, the server sends the user-edited script to a virtual character, integrating sound generation and visual presentation to run the event. The server uses Google's Text-to-Speech or Amazon Polly for speech synthesis software, and real-time 3D platforms such as Unity or Unreal Engine for visual effects. This allows the virtual character to behave like a real presenter, providing an engaging experience for the audience.

[0273] As described above, this invention combines various components to achieve efficient and flexible event management.

[0274] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0275] Step 1:

[0276] Users enter basic event information using the device's intuitive input interface. This includes the event title, date, participant list, agenda, and language selection. The entered information is stored on the device as digital data.

[0277] Step 2:

[0278] The terminal sends the event information received from the user to the server. The information is transferred using a secure communication protocol (e.g., HTTPS), and the server receives this data. The server analyzes the received data and performs a process to confirm the accuracy of the information.

[0279] Step 3:

[0280] Based on the received event information, the server uses a generative AI model to generate a progressive sentence. The language model performs natural language processing on the given prompt sentence and creates a script that progresses in an appropriate context. In this process, context is extracted from the input data (event information), and a sentence along the flow is output.

[0281] Step 4:

[0282] The server translates the generated progressive sentence into a plurality of specified languages using a multilingual conversion device. An automatic translation system is used for each language translation. The output is a translated script that includes culturally appropriate expressions adapted to different languages.

[0283] Step 5:

[0284] The server sends the translated script to the terminal. The terminal displays this script to the user, and the user can check and edit the content of the script as needed. Specific operations include modifying, adding, or deleting specific phrases. The edited script is saved on the terminal.

[0285] Step 6:

[0286] The server sends the final edited script to the virtual character. The virtual character uses dedicated audio generation software to synthesize voice and a system that displays visual expressions in real time. In this process, based on the script edited by the user, the event progresses in a natural form like a real host.

[0287] (Application Example 1)

[0288] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0289] Modern events increasingly require support for diverse cultural backgrounds and multilingualism, while simultaneously demanding the provision of realistic experiences in virtual environments. However, efficiently conveying information and managing participant interaction is extremely difficult. Similarly, in product introductions at virtual stores, effectively communicating the value of new products and maintaining participant interest is crucial. Traditional systems have faced challenges in adequately conveying information due to language barriers and cultural differences.

[0290] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0291] In this invention, the server includes means for providing an intuitive user interface for users to input basic event information, means for generating an event progress script using a natural language processing model based on the received information, and multilingual translation means for translating the generated script into multiple languages. This makes it possible to convey information to participants in a multilingual and culturally appropriate manner even in a virtual environment, and to provide participants with an immersive experience through product introductions in a virtual store.

[0292] A "user interface" is an intuitive and easy-to-use screen for users to input information into a system.

[0293] A "natural language processing model" is an algorithm that analyzes input text data to understand and generate human language.

[0294] An "event progress script" is a planned script of presentations designed to ensure an event runs smoothly.

[0295] A "multilingual translation method" is a function that accurately translates information across multiple different languages.

[0296] "Speech synthesis" is a technology that converts text information into speech that sounds like a human voice.

[0297] "Visual expression" refers to methods of conveying information visually through images and animations.

[0298] A "virtual presenter" is a software agent that acts as a digitally constructed moderator, providing information and event management support to the audience.

[0299] "Smart glasses" are wearable devices that can display virtual information in real time.

[0300] A "virtual store" is a digital store environment that is accessible to users online.

[0301] "Real-time information updates" is a feature that allows new information to be reflected instantly as the event progresses.

[0302] To realize this application, the system operates by integrating multiple technical components. The server is primarily responsible for data processing and generation. First, the user inputs basic event information into the terminal through an intuitive user interface. This interface is designed to be easy to operate and to quickly provide the user with the information they need.

[0303] The server uses the received information to utilize a natural language processing model and automatically generate an event progress script. This generated script is presented to the participants in real time within a virtual store using wearable devices such as smart glasses. The generated script is translated into multiple languages by a multilingual translation means. This enables appropriate and effective information transmission to international participants with different cultural backgrounds.

[0304] This script is presented by a virtual presenter that integrates speech synthesis technology and visual representation. This presenter provides a natural experience to the viewers and updates the information in real time according to the progress of the event. For example, when a user wearing smart glasses walks around a virtual store, a video introducing a new product is provided together with the voice. Also, it is possible to immediately generate relevant information in text using a generative AI model in response to the user's operations and questions.

[0305] As a specific example, a visual explaining the mechanism of a new product appears on the display of smart glasses. When the user requests detailed information with a voice command, the virtual presenter explains in detail the features and application examples of the product. An example of a generative AI prompt sentence is "Please write a script that explains the charm of the product from an angle that the user is likely to be interested in." In this way, the system can provide the participants with a deep sense of immersion and useful information.

[0306] The flow of the specific process in Application Example 1 will be described using FIG. 12.

[0307] Step 1:

[0308] The user uses the terminal to input the basic information of the event. The input information includes the title, schedule, participant list, and progress outline of the event. This information can be intuitively input via the user interface, and the terminal transmits this to the server.

[0309] Step 2:

[0310] The server uses natural language processing models to generate an event progress script based on the received event information. The input here is the event information, and the output is the initial draft of the script. This process involves analyzing the elements of the information and compiling the appropriate progress and important points into the script.

[0311] Step 3:

[0312] The generated script is translated into multiple languages ​​on the server by a multilingual translation engine. The input is the script in progress, and the output is the translated versions of the script. The translated script is adjusted to preserve the cultural nuances appropriate for each language.

[0313] Step 4:

[0314] Users can view the translated script on their device and edit its content using the interface. The input is the translated script, and the output is the final script edited by the user. The edited script is then sent back to the server.

[0315] Step 5:

[0316] The server uses the final script to integrate speech synthesis and visual representation, and controls the actions of the virtual presenter. The input is the final script, and the output is the virtual moderator's control commands for the presentation. This provides real-time narration and visual effects in accordance with the progress of the event.

[0317] Step 6:

[0318] Users wearing smart glasses are guided through a virtual store. Specific product information and explanations are presented in response to a script and real-time voice commands from the user. Input is user interaction, and output is product information presented as visual and audio data.

[0319] Step 7:

[0320] During the event, if any new information or changes arise, the server updates the script in real time and immediately reflects them in the virtual presenter. The input is real-time event information, and the output is the updated progress script. This step allows for flexible response to unexpected changes.

[0321] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0322] This invention provides a more advanced interactive experience by combining a virtual presenter system with an emotion engine that recognizes user emotions. The system includes a user interface, a natural language processing model, multilingual translation capabilities, a virtual presenter, integrated speech synthesis and visual representation functions, and an emotion engine. The specific operation flow is described below.

[0323] The user inputs event-related information through their device. This includes basic information as well as conditional settings depending on the situation. This information is sent to the server, which automatically generates an event progress script based on the information.

[0324] The generated script is translated into the required language through a multilingual translation system and adjusted to be a natural expression with cultural relevance. The translated script is then reviewed by the user on their device and edited and customized as needed.

[0325] A key feature here is the use of an emotion engine when the server prepares the virtual presenter's speech synthesis and visual presentation based on the edited script. The emotion engine analyzes the emotions of users and audience members in real time during the event. For example, if a user is nervous, the emotion engine recognizes this and adjusts the system to proceed in a more relaxed tone.

[0326] The analysis results from the emotion engine are transmitted to the server, which then dynamically adjusts the content of the script and the virtual presenter's expressions during the event. For example, if there are signs that the audience's interest is waning, the program will change the presenter's tone and gestures to try and make the presentation more engaging.

[0327] Ultimately, the terminal controls the event proceedings by a virtual presenter, using speech synthesis and visual representations to provide the audience with a realistic hosting experience. The system dynamically updates based on scripts and sentiment analysis results, allowing it to flexibly respond to unexpected situations during the event.

[0328] In this configuration, the system of the present invention can provide an interactive and personalized event experience that meets the needs of users and audiences.

[0329] The following describes the processing flow.

[0330] Step 1:

[0331] The user enters event details (e.g., event name, date, participant information, program details) using the user interface on their device. The device then formats the entered information appropriately and sends it to the server.

[0332] Step 2:

[0333] The server analyzes the information received from the user and generates an event progress script using a natural language processing model. The script includes introductions for each speaker and a description of the event flow in natural language.

[0334] Step 3:

[0335] The server translates the generated script into the required language via a multilingual translation system. The translated script is then adjusted to ensure culturally natural expression.

[0336] Step 4:

[0337] The device presents the translated script to the user. The user can review the script through the interface and edit or modify it as needed.

[0338] Step 5:

[0339] The server receives the edited script and prepares it for speech synthesis and visual representation. During this process, it activates the emotion engine and analyzes the emotions of the user or audience in real time.

[0340] Step 6:

[0341] The emotion engine analyzes the emotions of users and audience members and communicates the results to the server. Based on the information received, the server adjusts the script during the event and changes the tone and gestures of the virtual presenter.

[0342] Step 7:

[0343] During the event, the device controls the virtual presenter, using synthesized speech and animated visuals to manage the event. It also incorporates dynamic expressions to maintain audience engagement, taking into account analysis results from an emotion engine.

[0344] This series of processes allows the system to provide an interactive and personalized event experience and to flexibly adapt to any situation.

[0345] (Example 2)

[0346] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0347] Modern events are becoming increasingly diverse, requiring personalized experiences that meet the needs of participants. However, traditional methods have struggled to adapt to participants' emotions and cultural backgrounds in real time, making it difficult to create engaging and enjoyable events.

[0348] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0349] In this invention, the server includes means for providing an intuitive interface using a terminal for users to input information, means for generating a progress script using a data processing model based on the received information, and means for performing sentiment analysis and dynamically adjusting the facilitator's expressions according to the state of the audience. This makes it possible to provide a personalized interactive event experience that is tailored to the emotions and cultural backgrounds of the participants.

[0350] A "terminal" is an electronic device that allows a user to input information and interact with a system through an interface.

[0351] An "interface" is a screen or means available to a user for entering information into a system or for reviewing and editing scripts.

[0352] A "data processing model" is an algorithm or system that automatically generates event progress scripts based on received information.

[0353] A "procedure script" is a set of instructions that details the steps and content necessary to ensure an event runs smoothly.

[0354] "Multilingual translation means" refers to a technology or system used to automatically convert a generated progress script into multiple languages.

[0355] A "virtual presenter" is a computer-generated character that fulfills the role of a host or navigator, represented in a digital format.

[0356] "Speech generation and visual representation" refers to technologies that generate speech content as sound for a virtual presenter and provide visual effects for the audience.

[0357] "Emotion analysis" is the process of evaluating the emotional state of users or audiences in real time using sensors and algorithms.

[0358] This system includes terminals for user information input, a server that generates and controls progress scripts, and a mechanism for sentiment analysis. Users first input basic event information and participant characteristics via the terminal. The terminal provides an intuitive interface, making it easy for users to input information.

[0359] Information sent from the terminal is received by the server. The server uses a generative AI model to generate a natural language script based on the received information. The generated script is then translated into the required language using a multilingual translation tool. Tone and cultural appropriateness are also adjusted during this process.

[0360] For example, a possible prompt to input into a generative AI model is: "Generate a script for a 30-minute technical seminar. The target audience is young people, and it should include interactive elements."

[0361] The translated script is returned to the terminal, where the user can review and edit it as needed. Based on the user's edited script, the server adjusts the virtual presenter's voice generation and visual presentation, and prepares for the event. During this process, the server uses an emotion analysis engine to analyze the emotions of the user and audience in real time as the event progresses. Based on the emotion analysis results, the virtual presenter's expressions and tone are dynamically changed.

[0362] In this way, the system can provide users and audiences with a customized, interactive event experience. The system can also flexibly respond to unexpected situations during the event and make dynamic adjustments to ensure a realistic and engaging presentation.

[0363] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0364] Step 1:

[0365] Users use their devices to enter detailed information about the event. This information includes the date and time of the event, the theme, and the profile of the target audience. The entered information is organized as digital data on the device and sent to the server.

[0366] Step 2:

[0367] The server begins data processing based on the information received from the terminal. Specifically, it uses a generative AI model to generate a script for event progression. This model analyzes the input data based on the provided prompts and outputs natural language text. The generated script is then formatted as digital text data.

[0368] Step 3:

[0369] The server sends the generated script to a multilingual translation system. This system translates the script into the target language, evaluates and adjusts for cultural appropriateness. The translated script is then sent back to the terminal. During translation, language data conversion processing is performed, and the output is a natural expression in the target language.

[0370] Step 4:

[0371] The user reviews and edits the translated script on their device. If the user changes terminology in the script or adds new data, the device formats the edited data for resending to the server. This step involves text editing of the data.

[0372] Step 5:

[0373] The server prepares the virtual presenter based on the edited script. Specifically, it uses a speech generation engine to convert the script into audio data and a visual representation system to generate a visual presentation. In parallel, it uses an emotion analysis engine to analyze the emotions of users and the audience and generates data for real-time adjustments.

[0374] Step 6:

[0375] The terminal controls the event host in real time, receiving updated data from the server and dynamically adjusting the presentation. The system combines audio and visuals to create a realistic and interactive experience for the audience. Processing is performed to change the host's tone and gestures in response to the event situation and user feedback.

[0376] (Application Example 2)

[0377] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0378] In modern consumer activity, users require information about products and services that is tailored to their emotions and interests at that moment. However, conventional systems often lack sufficient real-time sentiment analysis and dynamic adjustment of information delivery, limiting the user experience. Furthermore, efficiently and culturally appropriate information delivery to global consumers who require multilingual support remains a challenge.

[0379] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0380] In this invention, the server includes means for providing an intuitive data input interface for users to input basic event information; means for generating an event progress script using a language processing device based on the received information; a translation device for translating the generated script into multiple languages; means for presenting the translated script to the user and providing a user-editable interface; means for controlling a virtual presenter by integrating speech synthesis and visual representation based on the edited script; an emotion analysis device for analyzing customer emotions and dynamically adjusting the information provided; and means for reflecting real-time information updates during the event. This enables personalized information provision in response to user emotions, thereby improving the user experience.

[0381] A "user" is someone who seeks information about products and services during their consumption activities.

[0382] A "data entry interface" is a means for users to intuitively input event and product information.

[0383] A "language processing device" is a device that analyzes and processes natural language and generates event progress scripts in various languages ​​based on the underlying information.

[0384] A "translation device" is a device used to translate a generated script into multiple languages.

[0385] An "interface" is a means by which the user can view and edit the provided translation script.

[0386] "Speech synthesis" is a technology that generates speech based on a processed script.

[0387] "Visual representation" refers to a means of visually representing a virtual narrator in conjunction with speech synthesis.

[0388] A "virtual facilitator" is a virtual display character that plays the role of guiding an event in a digital environment.

[0389] An "emotion analysis device" is a device that analyzes a user's emotions and dynamically adjusts the way information is provided based on that data.

[0390] "Personalization" refers to providing customized information tailored to the individual user's interests and emotions.

[0391] "Real-time" refers to something that happens immediately in the present moment.

[0392] To implement this invention, a server plays a central role. This server operates in conjunction with various software, as described below. First, the user inputs basic event information from a terminal through a data input interface. The terminal used here could be a general-purpose computer or smart device. Next, the server uses a language processing unit to perform natural language processing based on the input information and generates an event progress script. In this process, the server utilizes an advanced generative AI model.

[0393] The generated scripts are translated into multiple languages ​​using a translation tool. Services such as the Google Cloud Translation API can be used for this translation. The translated scripts are presented to the user and can be edited via a dedicated interface. This editing process also takes place on the user's device.

[0394] Once editing is complete, the server combines speech synthesis and visual representation technologies to generate a virtual presenter. Software such as Amazon Polly is used for speech synthesis, and the virtual presenter provides information to the user or audience. Furthermore, the server uses an emotion analyzer to analyze the emotions of the user or audience in real time. This analysis is then processed using machine learning libraries such as TensorFlow.

[0395] Based on the analysis results, the server dynamically adjusts how information is delivered to provide personalized information. For example, if the sentiment analyzer determines that the user seems uninterested while receiving a product description, the system will provide additional interesting information.

[0396] An example of a prompt message could be input to the generating AI model, such as, "Provide more interesting information based on the user's emotions." In this way, the present invention utilizes real-time user feedback to achieve effective information transmission.

[0397] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0398] Step 1:

[0399] The user enters basic event information into a data input interface via their device. The entered information is sent to the server as text data. The server receives this information and prepares it for natural language processing.

[0400] Step 2:

[0401] The server generates an event progress script using a language processing unit based on the received information. In this process, it utilizes a generative AI model to perform advanced natural language processing. The input is basic information provided by the user, and the output is the text data of the integrated event progress script.

[0402] Step 3:

[0403] The server uses a translation tool to translate the generated script into the required languages. The Google Cloud Translation API can be used for this purpose. The input is the text data of the generated script, and the output is the multilingual text data of the translated script.

[0404] Step 4:

[0405] The translated script is displayed on the user's terminal and becomes editable. The user reviews the information through the presented interface and makes corrections as needed. The input is the translated script, and the output is the script edited by the user.

[0406] Step 5:

[0407] Based on the edited script, the server integrates speech synthesis and visual representation to generate a virtual narrator. Amazon Polly or similar tools are used for speech synthesis, and the generated virtual narrator is displayed on the terminal. The input is the user-edited script, and the output is the virtual narrator's visual and audio data.

[0408] Step 6:

[0409] The server uses an emotion analysis device to analyze the emotions of users and audiences in real time. This analysis uses data acquired from the camera and microphone on the device, and is performed using TensorFlow or similar software. The input is the acquired emotion data, and the output is the emotion analysis result.

[0410] Step 7:

[0411] Based on the analysis results, the server dynamically adjusts how information is delivered. Specifically, the information and tone are changed according to the user's emotions. For example, the prompt "Please provide additional information to maintain the user's interest" is input into the generating AI model, and the customized information output is presented by a virtual presenter.

[0412] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0413] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0414] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0415] [Third Embodiment]

[0416] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0417] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0418] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0419] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0420] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0421] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0422] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0423] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0424] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0425] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0426] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0427] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0428] This invention provides a virtual presenter system for efficient and flexible event hosting. The system includes a user interface, a natural language processing model, multilingual translation capabilities, a virtual presenter, and integrated speech synthesis and visual representation functions. The specific operation flow is described below.

[0429] The user acts as the event organizer, using the terminal's user interface to input information related to the event. As a concrete example, consider information about an international conference. In this case, the user would input the event title, date, participant list, and a brief overview of the proceedings.

[0430] The entered information is sent to the server. The server receives this information and uses a natural language processing model to generate an event progress script. The generated script includes the flow of the moderator's speech and what to say, enabling smooth guidance during the event.

[0431] The server translates the progress script into multiple languages ​​as needed, enabling it to be accessible to international audiences. The translated script is then adjusted for culturally natural expression.

[0432] The translated and adjusted script is presented to the user on their device. The user can review this script and edit or customize specific parts as needed. For example, they can edit the details when introducing a particular speaker.

[0433] Finally, a virtual presenter guides the event using speech synthesis and visuals based on an edited script. The voice and gestures are synchronized at each stage, providing the audience with a realistic, presenter-like experience. The system can also accommodate real-time schedule changes during the event.

[0434] This configuration enables efficient, multilingual, and cost-effective moderation of events. This implementation is widely applicable to various types of events, and is particularly effective in international and multicultural settings.

[0435] The following describes the processing flow.

[0436] Step 1:

[0437] Users enter the necessary information for the event using the user interface on their device. This information includes the event name, date and time, participant information, and schedule.

[0438] Step 2:

[0439] The terminal sends information entered by the user to the server. The transmitted data is received by the server for processing.

[0440] Step 3:

[0441] The server analyzes the information it receives and automatically generates an event progress script using a natural language processing model. The script is created based on the input information and includes utterances and comments.

[0442] Step 4:

[0443] The server-generated scripts are fed into a translation system for multilingual support. The scripts are translated into the required languages ​​and culturally appropriate, with adjustments made to each language version.

[0444] Step 5:

[0445] The device presents the translated script to the user, who then reviews the content. The user edits and customizes the script for any parts that need correction.

[0446] Step 6:

[0447] The server receives the user's edited script and begins integrating the speech synthesis system and visual representation. It then sets up the virtual presenter's presentation, combining voice, facial expressions, and gestures.

[0448] Step 7:

[0449] On the day of the event, the device will begin the event proceedings with virtual presenters. It will output audio and animations in real time according to the script, and immediately reflect updates to information as needed.

[0450] This series of steps allows the system to manage events efficiently and flexibly.

[0451] (Example 1)

[0452] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0453] Efficiently and multilingually managing international and multicultural events is challenging. Furthermore, events with frequent schedule changes require rapid updates to the program. Existing technologies often require significant human resources and time to address these issues, resulting in high costs.

[0454] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0455] In this invention, the server includes means for providing an intuitive input device for users to input information, means for generating a progress statement using a language model based on the received information, and means for providing a multilingual conversion device for converting the generated statement into multiple languages. This enables smooth hosting in multiple languages ​​and allows for flexible response to information updates during the event.

[0456] A "user" is someone who operates the system, inputs information related to an event, and takes on the role of facilitating its progress.

[0457] An "input device" refers to a device or interface that a user uses to input information.

[0458] A "language model" is a natural language processing technique used to generate a continuous sentence based on input information.

[0459] A "proceeding document" refers to the script or record used within an event, and contains information necessary for the smooth running of the event.

[0460] A "translation device" refers to a technology or means used to translate a generated text into multiple different languages.

[0461] "Multilingual support" means the ability to support multiple languages.

[0462] A "virtual character" is a digital character that integrates sound generation and visual expression to conduct events on screen.

[0463] "Information updates" refer to the immediate reflection of any changes or new information that arise as the event progresses.

[0464] This invention realizes a virtual MC system for efficiently managing events, in which a server, terminal, and user work together.

[0465] First, the user uses the device's intuitive input device to enter information related to the event. Specifically, they can enter the event title, date, participant information, and a summary of the agenda. An example of a prompt message would be: "Please create a progress script for the next event. The title is 'International Conference on Technological Innovation,' the date is November 15, 2023, and the participants are experts from around the world."

[0466] The terminal sends this input information to the server. The server uses a language model based on the received information to generate a continuous sentence. This language model is based on natural language processing technology, and it is conceivable that existing technologies such as Google's TensorFlow or OpenAI's GPT could be used.

[0467] The server translates the generated proceedings into multiple languages. Machine translation systems are used for each language, with platforms such as Microsoft Translator and Amazon Translate available. This multilingual support ensures smooth proceedings for audiences at international events.

[0468] The translated script is sent to the user's device, where they can review its content and make adjustments or edits as needed. For example, it's easy to add or modify names and job titles when introducing specific speakers.

[0469] Ultimately, the server sends the user-edited script to a virtual character, integrating sound generation and visual presentation to run the event. The server uses Google's Text-to-Speech or Amazon Polly for speech synthesis software, and real-time 3D platforms such as Unity or Unreal Engine for visual effects. This allows the virtual character to behave like a real presenter, providing an engaging experience for the audience.

[0470] As described above, this invention combines various components to achieve efficient and flexible event management.

[0471] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0472] Step 1:

[0473] Users enter basic event information using the device's intuitive input interface. This includes the event title, date, participant list, agenda, and language selection. The entered information is stored on the device as digital data.

[0474] Step 2:

[0475] The terminal sends event information received from the user to the server. The information is transferred using a secure communication protocol (e.g., HTTPS), and the server receives this data. The server analyzes the received data and verifies its accuracy.

[0476] Step 3:

[0477] The server generates a sequence of events using a generative AI model based on the received event information. The language model performs natural language processing on the given prompt sentence and creates a script that progresses within the appropriate context. In this process, context is extracted from the input data (event information), and sentences that follow the flow are output.

[0478] Step 4:

[0479] The server translates the generated text into multiple specified languages ​​using a multilingual translation device. An automatic translation system is used for each language. The output is a translated script containing culturally appropriate expressions for each language.

[0480] Step 5:

[0481] The server sends the translated script to the terminal. The terminal displays this script to the user, who can review and edit its contents as needed. Specific actions include modifying, adding, or deleting specific phrases. The edited script is saved to the terminal.

[0482] Step 6:

[0483] The server sends the final edited script to the virtual character. The virtual character uses specialized sound generation software to synthesize speech and a system that displays visual representations in real time. In this process, the virtual character conducts the event in a natural manner, like a real presenter, based on the script edited by the user.

[0484] (Application Example 1)

[0485] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0486] Modern events increasingly require support for diverse cultural backgrounds and multilingualism, while simultaneously demanding the provision of realistic experiences in virtual environments. However, efficiently conveying information and managing participant interaction is extremely difficult. Similarly, in product introductions at virtual stores, effectively communicating the value of new products and maintaining participant interest is crucial. Traditional systems have faced challenges in adequately conveying information due to language barriers and cultural differences.

[0487] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0488] In this invention, the server includes means for providing an intuitive user interface for users to input basic event information, means for generating an event progress script using a natural language processing model based on the received information, and multilingual translation means for translating the generated script into multiple languages. This makes it possible to convey information to participants in a multilingual and culturally appropriate manner even in a virtual environment, and to provide participants with an immersive experience through product introductions in a virtual store.

[0489] A "user interface" is an intuitive and easy-to-use screen for users to input information into a system.

[0490] A "natural language processing model" is an algorithm that analyzes input text data to understand and generate human language.

[0491] An "event progress script" is a planned script of presentations designed to ensure an event runs smoothly.

[0492] A "multilingual translation method" is a function that accurately translates information across multiple different languages.

[0493] "Speech synthesis" is a technology that converts text information into speech that sounds like a human voice.

[0494] "Visual expression" refers to methods of conveying information visually through images and animations.

[0495] A "virtual presenter" is a software agent that acts as a digitally constructed moderator, providing information and event management support to the audience.

[0496] "Smart glasses" are wearable devices that can display virtual information in real time.

[0497] A "virtual store" is a digital store environment that is accessible to users online.

[0498] "Real-time information updates" is a feature that allows new information to be reflected instantly as the event progresses.

[0499] To realize this application, the system operates by integrating multiple technical components. The server is primarily responsible for data processing and generation. First, the user inputs basic event information into the terminal through an intuitive user interface. This interface is designed to be easy to operate and to quickly provide the user with the information they need.

[0500] The server uses the received information to automatically generate an event progress script, utilizing a natural language processing model. This generated script is presented to participants in real time within the virtual store using wearable devices such as smart glasses. The generated script is translated into multiple languages ​​using a multilingual translation system. This enables appropriate and effective information transmission to international participants with diverse cultural backgrounds.

[0501] This script is presented by a virtual presenter that integrates speech synthesis technology and visual representation. This presenter provides a natural experience for the audience, updating information in real time as the event progresses. For example, as a user wearing smart glasses walks around a virtual store, it can provide introductory videos of new products along with audio. It can also instantly generate relevant information in text using a generative AI model in response to user actions and questions.

[0502] A concrete example is a scenario where a visual explaining the mechanism of a new product appears on the display of smart glasses, and the user requests more information using voice commands, at which point a virtual presenter explains the product's features and applications in detail. An example of a generated AI prompt is, "Please write a script that explains the product's appeal from an angle that the user is likely to find interesting." In this way, the system can provide participants with a deep sense of immersion and useful information.

[0503] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0504] Step 1:

[0505] Users enter basic event information using their devices. This information includes the event title, date, participant list, and agenda. This information can be entered intuitively through the user interface, and the device then sends it to the server.

[0506] Step 2:

[0507] The server uses natural language processing models to generate an event progress script based on the received event information. The input here is the event information, and the output is the initial draft of the script. This process involves analyzing the elements of the information and compiling the appropriate progress and important points into the script.

[0508] Step 3:

[0509] The generated script is translated into multiple languages ​​on the server by a multilingual translation engine. The input is the script in progress, and the output is the translated versions of the script. The translated script is adjusted to preserve the cultural nuances appropriate for each language.

[0510] Step 4:

[0511] Users can view the translated script on their device and edit its content using the interface. The input is the translated script, and the output is the final script edited by the user. The edited script is then sent back to the server.

[0512] Step 5:

[0513] The server uses the final script to integrate speech synthesis and visual representation, and controls the actions of the virtual presenter. The input is the final script, and the output is the virtual moderator's control commands for the presentation. This provides real-time narration and visual effects in accordance with the progress of the event.

[0514] Step 6:

[0515] Users wearing smart glasses are guided through a virtual store. Specific product information and explanations are presented in response to a script and real-time voice commands from the user. Input is user interaction, and output is product information presented as visual and audio data.

[0516] Step 7:

[0517] During the event, if any new information or changes arise, the server updates the script in real time and immediately reflects them in the virtual presenter. The input is real-time event information, and the output is the updated progress script. This step allows for flexible response to unexpected changes.

[0518] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0519] This invention provides a more advanced interactive experience by combining a virtual presenter system with an emotion engine that recognizes user emotions. The system includes a user interface, a natural language processing model, multilingual translation capabilities, a virtual presenter, integrated speech synthesis and visual representation functions, and an emotion engine. The specific operation flow is described below.

[0520] The user inputs event-related information through their device. This includes basic information as well as conditional settings depending on the situation. This information is sent to the server, which automatically generates an event progress script based on the information.

[0521] The generated script is translated into the required language through a multilingual translation system and adjusted to be a natural expression with cultural relevance. The translated script is then reviewed by the user on their device and edited and customized as needed.

[0522] A key feature here is the use of an emotion engine when the server prepares the virtual presenter's speech synthesis and visual presentation based on the edited script. The emotion engine analyzes the emotions of users and audience members in real time during the event. For example, if a user is nervous, the emotion engine recognizes this and adjusts the system to proceed in a more relaxed tone.

[0523] The analysis results from the emotion engine are transmitted to the server, which then dynamically adjusts the content of the script and the virtual presenter's expressions during the event. For example, if there are signs that the audience's interest is waning, the program will change the presenter's tone and gestures to try and make the presentation more engaging.

[0524] Ultimately, the terminal controls the event proceedings by a virtual presenter, using speech synthesis and visual representations to provide the audience with a realistic hosting experience. The system dynamically updates based on scripts and sentiment analysis results, allowing it to flexibly respond to unexpected situations during the event.

[0525] In this configuration, the system of the present invention can provide an interactive and personalized event experience that meets the needs of users and audiences.

[0526] The following describes the processing flow.

[0527] Step 1:

[0528] The user enters event details (e.g., event name, date, participant information, program details) using the user interface on their device. The device then formats the entered information appropriately and sends it to the server.

[0529] Step 2:

[0530] The server analyzes the information received from the user and generates an event progress script using a natural language processing model. The script includes introductions for each speaker and a description of the event flow in natural language.

[0531] Step 3:

[0532] The server translates the generated script into the required language via a multilingual translation system. The translated script is then adjusted to ensure culturally natural expression.

[0533] Step 4:

[0534] The device presents the translated script to the user. The user can review the script through the interface and edit or modify it as needed.

[0535] Step 5:

[0536] The server receives the edited script and prepares it for speech synthesis and visual representation. During this process, it activates the emotion engine and analyzes the emotions of the user or audience in real time.

[0537] Step 6:

[0538] The emotion engine analyzes the emotions of users and audience members and communicates the results to the server. Based on the information received, the server adjusts the script during the event and changes the tone and gestures of the virtual presenter.

[0539] Step 7:

[0540] During the event, the device controls the virtual presenter, using synthesized speech and animated visuals to manage the event. It also incorporates dynamic expressions to maintain audience engagement, taking into account analysis results from an emotion engine.

[0541] This series of processes allows the system to provide an interactive and personalized event experience and to flexibly adapt to any situation.

[0542] (Example 2)

[0543] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0544] Modern events are becoming increasingly diverse, requiring personalized experiences that meet the needs of participants. However, traditional methods have struggled to adapt to participants' emotions and cultural backgrounds in real time, making it difficult to create engaging and enjoyable events.

[0545] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0546] In this invention, the server includes means for providing an intuitive interface using a terminal for users to input information, means for generating a progress script using a data processing model based on the received information, and means for performing sentiment analysis and dynamically adjusting the facilitator's expressions according to the state of the audience. This makes it possible to provide a personalized interactive event experience that is tailored to the emotions and cultural backgrounds of the participants.

[0547] A "terminal" is an electronic device that allows a user to input information and interact with a system through an interface.

[0548] An "interface" is a screen or means available to a user for entering information into a system or for reviewing and editing scripts.

[0549] A "data processing model" is an algorithm or system that automatically generates event progress scripts based on received information.

[0550] A "procedure script" is a set of instructions that details the steps and content necessary to ensure an event runs smoothly.

[0551] "Multilingual translation means" refers to a technology or system used to automatically convert a generated progress script into multiple languages.

[0552] A "virtual presenter" is a computer-generated character that fulfills the role of a host or navigator, represented in a digital format.

[0553] "Speech generation and visual representation" refers to technologies that generate speech content as sound for a virtual presenter and provide visual effects for the audience.

[0554] "Emotion analysis" is the process of evaluating the emotional state of users or audiences in real time using sensors and algorithms.

[0555] This system includes terminals for user information input, a server that generates and controls progress scripts, and a mechanism for sentiment analysis. Users first input basic event information and participant characteristics via the terminal. The terminal provides an intuitive interface, making it easy for users to input information.

[0556] Information sent from the terminal is received by the server. The server uses a generative AI model to generate a natural language script based on the received information. The generated script is then translated into the required language using a multilingual translation tool. Tone and cultural appropriateness are also adjusted during this process.

[0557] For example, a possible prompt to input into a generative AI model is: "Generate a script for a 30-minute technical seminar. The target audience is young people, and it should include interactive elements."

[0558] The translated script is returned to the terminal, where the user can review and edit it as needed. Based on the user's edited script, the server adjusts the virtual presenter's voice generation and visual presentation, and prepares for the event. During this process, the server uses an emotion analysis engine to analyze the emotions of the user and audience in real time as the event progresses. Based on the emotion analysis results, the virtual presenter's expressions and tone are dynamically changed.

[0559] In this way, the system can provide users and audiences with a customized, interactive event experience. The system can also flexibly respond to unexpected situations during the event and make dynamic adjustments to ensure a realistic and engaging presentation.

[0560] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0561] Step 1:

[0562] Users use their devices to enter detailed information about the event. This information includes the date and time of the event, the theme, and the profile of the target audience. The entered information is organized as digital data on the device and sent to the server.

[0563] Step 2:

[0564] The server begins data processing based on the information received from the terminal. Specifically, it uses a generative AI model to generate a script for event progression. This model analyzes the input data based on the provided prompts and outputs natural language text. The generated script is then formatted as digital text data.

[0565] Step 3:

[0566] The server sends the generated script to a multilingual translation system. This system translates the script into the target language, evaluates and adjusts for cultural appropriateness. The translated script is then sent back to the terminal. During translation, language data conversion processing is performed, and the output is a natural expression in the target language.

[0567] Step 4:

[0568] The user reviews and edits the translated script on their device. If the user changes terminology in the script or adds new data, the device formats the edited data for resending to the server. This step involves text editing of the data.

[0569] Step 5:

[0570] The server prepares the virtual presenter based on the edited script. Specifically, it uses a speech generation engine to convert the script into audio data and a visual representation system to generate a visual presentation. In parallel, it uses an emotion analysis engine to analyze the emotions of users and the audience and generates data for real-time adjustments.

[0571] Step 6:

[0572] The terminal controls the event host in real time, receiving updated data from the server and dynamically adjusting the presentation. The system combines audio and visuals to create a realistic and interactive experience for the audience. Processing is performed to change the host's tone and gestures in response to the event situation and user feedback.

[0573] (Application Example 2)

[0574] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0575] In modern consumer activity, users require information about products and services that is tailored to their emotions and interests at that moment. However, conventional systems often lack sufficient real-time sentiment analysis and dynamic adjustment of information delivery, limiting the user experience. Furthermore, efficiently and culturally appropriate information delivery to global consumers who require multilingual support remains a challenge.

[0576] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0577] In this invention, the server includes means for providing an intuitive data input interface for users to input basic event information; means for generating an event progress script using a language processing device based on the received information; a translation device for translating the generated script into multiple languages; means for presenting the translated script to the user and providing a user-editable interface; means for controlling a virtual presenter by integrating speech synthesis and visual representation based on the edited script; an emotion analysis device for analyzing customer emotions and dynamically adjusting the information provided; and means for reflecting real-time information updates during the event. This enables personalized information provision in response to user emotions, thereby improving the user experience.

[0578] A "user" is someone who seeks information about products and services during their consumption activities.

[0579] A "data entry interface" is a means for users to intuitively input event and product information.

[0580] A "language processing device" is a device that analyzes and processes natural language and generates event progress scripts in various languages ​​based on the underlying information.

[0581] A "translation device" is a device used to translate a generated script into multiple languages.

[0582] An "interface" is a means by which the user can view and edit the provided translation script.

[0583] "Speech synthesis" is a technology that generates speech based on a processed script.

[0584] "Visual representation" refers to a means of visually representing a virtual narrator in conjunction with speech synthesis.

[0585] A "virtual facilitator" is a virtual display character that plays the role of guiding an event in a digital environment.

[0586] An "emotion analysis device" is a device that analyzes a user's emotions and dynamically adjusts the way information is provided based on that data.

[0587] "Personalization" refers to providing customized information tailored to the individual user's interests and emotions.

[0588] "Real-time" refers to something that happens immediately in the present moment.

[0589] To implement this invention, a server plays a central role. This server operates in conjunction with various software, as described below. First, the user inputs basic event information from a terminal through a data input interface. The terminal used here could be a general-purpose computer or smart device. Next, the server uses a language processing unit to perform natural language processing based on the input information and generates an event progress script. In this process, the server utilizes an advanced generative AI model.

[0590] The generated scripts are translated into multiple languages ​​using a translation tool. Services such as the Google Cloud Translation API can be used for this translation. The translated scripts are presented to the user and can be edited via a dedicated interface. This editing process also takes place on the user's device.

[0591] Once editing is complete, the server combines speech synthesis and visual representation technologies to generate a virtual presenter. Software such as Amazon Polly is used for speech synthesis, and the virtual presenter provides information to the user or audience. Furthermore, the server uses an emotion analyzer to analyze the emotions of the user or audience in real time. This analysis is then processed using machine learning libraries such as TensorFlow.

[0592] Based on the analysis results, the server dynamically adjusts how information is delivered to provide personalized information. For example, if the sentiment analyzer determines that the user seems uninterested while receiving a product description, the system will provide additional interesting information.

[0593] An example of a prompt message could be input to the generating AI model, such as, "Provide more interesting information based on the user's emotions." In this way, the present invention utilizes real-time user feedback to achieve effective information transmission.

[0594] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0595] Step 1:

[0596] The user enters basic event information into a data input interface via their device. The entered information is sent to the server as text data. The server receives this information and prepares it for natural language processing.

[0597] Step 2:

[0598] The server generates an event progress script using a language processing unit based on the received information. In this process, it utilizes a generative AI model to perform advanced natural language processing. The input is basic information provided by the user, and the output is the text data of the integrated event progress script.

[0599] Step 3:

[0600] The server uses a translation tool to translate the generated script into the required languages. The Google Cloud Translation API can be used for this purpose. The input is the text data of the generated script, and the output is the multilingual text data of the translated script.

[0601] Step 4:

[0602] The translated script is displayed on the user's terminal and becomes editable. The user reviews the information through the presented interface and makes corrections as needed. The input is the translated script, and the output is the script edited by the user.

[0603] Step 5:

[0604] Based on the edited script, the server integrates speech synthesis and visual representation to generate a virtual narrator. Amazon Polly or similar tools are used for speech synthesis, and the generated virtual narrator is displayed on the terminal. The input is the user-edited script, and the output is the virtual narrator's visual and audio data.

[0605] Step 6:

[0606] The server uses an emotion analysis device to analyze the emotions of users and audiences in real time. This analysis uses data acquired from the camera and microphone on the device, and is performed using TensorFlow or similar software. The input is the acquired emotion data, and the output is the emotion analysis result.

[0607] Step 7:

[0608] Based on the analysis results, the server dynamically adjusts how information is delivered. Specifically, the information and tone are changed according to the user's emotions. For example, the prompt "Please provide additional information to maintain the user's interest" is input into the generating AI model, and the customized information output is presented by a virtual presenter.

[0609] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0610] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0611] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0612] [Fourth Embodiment]

[0613] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0614] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0615] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0616] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0617] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0618] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0619] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0620] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0621] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0622] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0623] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0624] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0625] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0626] This invention provides a virtual presenter system for efficient and flexible event hosting. The system includes a user interface, a natural language processing model, multilingual translation capabilities, a virtual presenter, and integrated speech synthesis and visual representation functions. The specific operation flow is described below.

[0627] The user acts as the event organizer, using the terminal's user interface to input information related to the event. As a concrete example, consider information about an international conference. In this case, the user would input the event title, date, participant list, and a brief overview of the proceedings.

[0628] The entered information is sent to the server. The server receives this information and uses a natural language processing model to generate an event progress script. The generated script includes the flow of the moderator's speech and what to say, enabling smooth guidance during the event.

[0629] The server translates the progress script into multiple languages ​​as needed, enabling it to be accessible to international audiences. The translated script is then adjusted for culturally natural expression.

[0630] The translated and adjusted script is presented to the user on their device. The user can review this script and edit or customize specific parts as needed. For example, they can edit the details when introducing a particular speaker.

[0631] Finally, a virtual presenter guides the event using speech synthesis and visuals based on an edited script. The voice and gestures are synchronized at each stage, providing the audience with a realistic, presenter-like experience. The system can also accommodate real-time schedule changes during the event.

[0632] This configuration enables efficient, multilingual, and cost-effective moderation of events. This implementation is widely applicable to various types of events, and is particularly effective in international and multicultural settings.

[0633] The following describes the processing flow.

[0634] Step 1:

[0635] Users enter the necessary information for the event using the user interface on their device. This information includes the event name, date and time, participant information, and schedule.

[0636] Step 2:

[0637] The terminal sends information entered by the user to the server. The transmitted data is received by the server for processing.

[0638] Step 3:

[0639] The server analyzes the information it receives and automatically generates an event progress script using a natural language processing model. The script is created based on the input information and includes utterances and comments.

[0640] Step 4:

[0641] The server-generated scripts are fed into a translation system for multilingual support. The scripts are translated into the required languages ​​and culturally appropriate, with adjustments made to each language version.

[0642] Step 5:

[0643] The device presents the translated script to the user, who then reviews the content. The user edits and customizes the script for any parts that need correction.

[0644] Step 6:

[0645] The server receives the user's edited script and begins integrating the speech synthesis system and visual representation. It then sets up the virtual presenter's presentation, combining voice, facial expressions, and gestures.

[0646] Step 7:

[0647] On the day of the event, the device will begin the event proceedings with virtual presenters. It will output audio and animations in real time according to the script, and immediately reflect updates to information as needed.

[0648] This series of steps allows the system to manage events efficiently and flexibly.

[0649] (Example 1)

[0650] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0651] Efficiently and multilingually managing international and multicultural events is challenging. Furthermore, events with frequent schedule changes require rapid updates to the program. Existing technologies often require significant human resources and time to address these issues, resulting in high costs.

[0652] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0653] In this invention, the server includes means for providing an intuitive input device for users to input information, means for generating a progress statement using a language model based on the received information, and means for providing a multilingual conversion device for converting the generated statement into multiple languages. This enables smooth hosting in multiple languages ​​and allows for flexible response to information updates during the event.

[0654] A "user" is someone who operates the system, inputs information related to an event, and takes on the role of facilitating its progress.

[0655] An "input device" refers to a device or interface that a user uses to input information.

[0656] A "language model" is a natural language processing technique used to generate a continuous sentence based on input information.

[0657] A "proceeding document" refers to the script or record used within an event, and contains information necessary for the smooth running of the event.

[0658] A "translation device" refers to a technology or means used to translate a generated text into multiple different languages.

[0659] "Multilingual support" means the ability to support multiple languages.

[0660] A "virtual character" is a digital character that integrates sound generation and visual expression to conduct events on screen.

[0661] "Information updates" refer to the immediate reflection of any changes or new information that arise as the event progresses.

[0662] This invention realizes a virtual MC system for efficiently managing events, in which a server, terminal, and user work together.

[0663] First, the user uses the device's intuitive input device to enter information related to the event. Specifically, they can enter the event title, date, participant information, and a summary of the agenda. An example of a prompt message would be: "Please create a progress script for the next event. The title is 'International Conference on Technological Innovation,' the date is November 15, 2023, and the participants are experts from around the world."

[0664] The terminal sends this input information to the server. The server uses a language model based on the received information to generate a continuous sentence. This language model is based on natural language processing technology, and it is conceivable that existing technologies such as Google's TensorFlow or OpenAI's GPT could be used.

[0665] The server translates the generated proceedings into multiple languages. Machine translation systems are used for each language, with platforms such as Microsoft Translator and Amazon Translate available. This multilingual support ensures smooth proceedings for audiences at international events.

[0666] The translated script is sent to the device, where the user can review its content and make adjustments or edits as needed. For example, it is easy to add or modify the names and job titles of specific speakers when introducing them.

[0667] Ultimately, the server sends the user-edited script to a virtual character, integrating sound generation and visual presentation to run the event. The server uses Google's Text-to-Speech or Amazon Polly for speech synthesis software, and real-time 3D platforms such as Unity or Unreal Engine for visual effects. This allows the virtual character to behave like a real presenter, providing an engaging experience for the audience.

[0668] As described above, this invention combines various components to achieve efficient and flexible event management.

[0669] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0670] Step 1:

[0671] Users enter basic event information using the device's intuitive input interface. This includes the event title, date, participant list, agenda, and language selection. The entered information is stored on the device as digital data.

[0672] Step 2:

[0673] The terminal sends event information received from the user to the server. The information is transferred using a secure communication protocol (e.g., HTTPS), and the server receives this data. The server analyzes the received data and verifies its accuracy.

[0674] Step 3:

[0675] The server generates a sequence of events using a generative AI model based on the received event information. The language model performs natural language processing on the given prompt sentence and creates a script that progresses within the appropriate context. In this process, context is extracted from the input data (event information), and sentences that follow the flow are output.

[0676] Step 4:

[0677] The server translates the generated text into multiple specified languages ​​using a multilingual translation device. An automatic translation system is used for each language. The output is a translated script containing culturally appropriate expressions for each language.

[0678] Step 5:

[0679] The server sends the translated script to the terminal. The terminal displays this script to the user, who can review and edit its contents as needed. Specific actions include modifying, adding, or deleting specific phrases. The edited script is saved to the terminal.

[0680] Step 6:

[0681] The server sends the final edited script to the virtual character. The virtual character uses specialized sound generation software to synthesize speech and a system that displays visual representations in real time. In this process, the virtual character conducts the event in a natural manner, like a real presenter, based on the script edited by the user.

[0682] (Application Example 1)

[0683] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0684] Modern events increasingly require support for diverse cultural backgrounds and multilingualism, while simultaneously demanding the provision of realistic experiences in virtual environments. However, efficiently conveying information and managing participant interaction is extremely difficult. Similarly, in product introductions at virtual stores, effectively communicating the value of new products and maintaining participant interest is crucial. Traditional systems have faced challenges in adequately conveying information due to language barriers and cultural differences.

[0685] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0686] In this invention, the server includes means for providing an intuitive user interface for users to input basic event information, means for generating an event progress script using a natural language processing model based on the received information, and multilingual translation means for translating the generated script into multiple languages. This makes it possible to convey information to participants in a multilingual and culturally appropriate manner even in a virtual environment, and to provide participants with an immersive experience through product introductions in a virtual store.

[0687] A "user interface" is an intuitive and easy-to-use screen for users to input information into a system.

[0688] A "natural language processing model" is an algorithm that analyzes input text data to understand and generate human language.

[0689] An "event progress script" is a planned script of presentations designed to ensure an event runs smoothly.

[0690] A "multilingual translation method" is a function that accurately translates information across multiple different languages.

[0691] "Speech synthesis" is a technology that converts text information into speech that sounds like a human voice.

[0692] "Visual expression" refers to methods of conveying information visually through images and animations.

[0693] A "virtual presenter" is a software agent that acts as a digitally constructed moderator, providing information and event management support to the audience.

[0694] "Smart glasses" are wearable devices that can display virtual information in real time.

[0695] A "virtual store" is a digital store environment that is accessible to users online.

[0696] "Real-time information updates" is a feature that allows new information to be reflected instantly as the event progresses.

[0697] To realize this application, the system operates by integrating multiple technical components. The server is primarily responsible for data processing and generation. First, the user inputs basic event information into the terminal through an intuitive user interface. This interface is designed to be easy to operate and to quickly provide the user with the information they need.

[0698] The server uses the received information to automatically generate an event progress script, utilizing a natural language processing model. This generated script is presented to participants in real time within the virtual store using wearable devices such as smart glasses. The generated script is translated into multiple languages ​​using a multilingual translation system. This enables appropriate and effective information transmission to international participants with diverse cultural backgrounds.

[0699] This script is presented by a virtual presenter that integrates speech synthesis technology and visual representation. This presenter provides a natural experience for the audience, updating information in real time as the event progresses. For example, as a user wearing smart glasses walks around a virtual store, it can provide introductory videos of new products along with audio. It can also instantly generate relevant information in text using a generative AI model in response to user actions and questions.

[0700] A concrete example is a scenario where a visual explaining the mechanism of a new product appears on the display of smart glasses, and the user requests more information using voice commands, at which point a virtual presenter explains the product's features and applications in detail. An example of a generated AI prompt is, "Please write a script that explains the product's appeal from an angle that the user is likely to find interesting." In this way, the system can provide participants with a deep sense of immersion and useful information.

[0701] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0702] Step 1:

[0703] Users enter basic event information using their devices. This information includes the event title, date, participant list, and agenda. This information can be entered intuitively through the user interface, and the device then sends it to the server.

[0704] Step 2:

[0705] The server uses natural language processing models to generate an event progress script based on the received event information. The input here is the event information, and the output is the initial draft of the script. This process involves analyzing the elements of the information and compiling the appropriate progress and important points into the script.

[0706] Step 3:

[0707] The generated script is translated into multiple languages ​​on the server by a multilingual translation engine. The input is the script in progress, and the output is the translated versions of the script. The translated script is adjusted to preserve the cultural nuances appropriate for each language.

[0708] Step 4:

[0709] Users can view the translated script on their device and edit its content using the interface. The input is the translated script, and the output is the final script edited by the user. The edited script is then sent back to the server.

[0710] Step 5:

[0711] The server uses the final script to integrate speech synthesis and visual representation, and controls the actions of the virtual presenter. The input is the final script, and the output is the virtual moderator's control commands for the presentation. This provides real-time narration and visual effects in accordance with the progress of the event.

[0712] Step 6:

[0713] Users wearing smart glasses are guided through a virtual store. Specific product information and explanations are presented in response to a script and real-time voice commands from the user. Input is user interaction, and output is product information presented as visual and audio data.

[0714] Step 7:

[0715] During the event, if any new information or changes arise, the server updates the script in real time and immediately reflects them in the virtual presenter. The input is real-time event information, and the output is the updated progress script. This step allows for flexible response to unexpected changes.

[0716] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0717] This invention provides a more advanced interactive experience by combining a virtual presenter system with an emotion engine that recognizes user emotions. The system includes a user interface, a natural language processing model, multilingual translation capabilities, a virtual presenter, integrated speech synthesis and visual representation functions, and an emotion engine. The specific operation flow is described below.

[0718] The user inputs event-related information through their device. This includes basic information as well as conditional settings depending on the situation. This information is sent to the server, which automatically generates an event progress script based on the information.

[0719] The generated script is translated into the required language through a multilingual translation system and adjusted to be a natural expression with cultural relevance. The translated script is then reviewed by the user on their device and edited and customized as needed.

[0720] A key feature here is the use of an emotion engine when the server prepares the virtual presenter's speech synthesis and visual presentation based on the edited script. The emotion engine analyzes the emotions of users and audience members in real time during the event. For example, if a user is nervous, the emotion engine recognizes this and adjusts the system to proceed in a more relaxed tone.

[0721] The analysis results from the emotion engine are transmitted to the server, which then dynamically adjusts the content of the script and the virtual presenter's expressions during the event. For example, if there are signs that the audience's interest is waning, the program will change the presenter's tone and gestures to try and make the presentation more engaging.

[0722] Ultimately, the terminal controls the event proceedings by a virtual presenter, using speech synthesis and visual representations to provide the audience with a realistic hosting experience. The system dynamically updates based on scripts and sentiment analysis results, allowing it to flexibly respond to unexpected situations during the event.

[0723] In this configuration, the system of the present invention can provide an interactive and personalized event experience that meets the needs of users and audiences.

[0724] The following describes the processing flow.

[0725] Step 1:

[0726] The user enters event details (e.g., event name, date, participant information, program details) using the user interface on their device. The device then formats the entered information appropriately and sends it to the server.

[0727] Step 2:

[0728] The server analyzes the information received from the user and generates an event progress script using a natural language processing model. The script includes introductions for each speaker and a description of the event flow in natural language.

[0729] Step 3:

[0730] The server translates the generated script into the required language via a multilingual translation system. The translated script is then adjusted to ensure culturally natural expression.

[0731] Step 4:

[0732] The device presents the translated script to the user. The user can review the script through the interface and edit or modify it as needed.

[0733] Step 5:

[0734] The server receives the edited script and prepares it for speech synthesis and visual representation. During this process, it activates the emotion engine and analyzes the emotions of the user or audience in real time.

[0735] Step 6:

[0736] The emotion engine analyzes the emotions of users and audience members and communicates the results to the server. Based on the information received, the server adjusts the script during the event and changes the tone and gestures of the virtual presenter.

[0737] Step 7:

[0738] During the event, the device controls the virtual presenter, using synthesized speech and animated visuals to manage the event. It also incorporates dynamic expressions to maintain audience engagement, taking into account analysis results from an emotion engine.

[0739] This series of processes allows the system to provide an interactive and personalized event experience and to flexibly adapt to any situation.

[0740] (Example 2)

[0741] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0742] Modern events are becoming increasingly diverse, requiring personalized experiences that meet the needs of participants. However, traditional methods have struggled to adapt to participants' emotions and cultural backgrounds in real time, making it difficult to create engaging and enjoyable events.

[0743] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0744] In this invention, the server includes means for providing an intuitive interface using a terminal for users to input information, means for generating a progress script using a data processing model based on the received information, and means for performing sentiment analysis and dynamically adjusting the facilitator's expressions according to the state of the audience. This makes it possible to provide a personalized interactive event experience that is tailored to the emotions and cultural backgrounds of the participants.

[0745] A "terminal" is an electronic device that allows a user to input information and interact with a system through an interface.

[0746] An "interface" is a screen or means available to a user for entering information into a system or for reviewing and editing scripts.

[0747] A "data processing model" is an algorithm or system that automatically generates event progress scripts based on received information.

[0748] A "procedure script" is a set of instructions that details the steps and content necessary to ensure an event runs smoothly.

[0749] "Multilingual translation means" refers to a technology or system used to automatically convert a generated progress script into multiple languages.

[0750] A "virtual presenter" is a computer-generated character that fulfills the role of a host or navigator, represented in a digital format.

[0751] "Speech generation and visual representation" refers to technologies that generate speech content as sound for a virtual presenter and provide visual effects for the audience.

[0752] "Emotion analysis" is the process of evaluating the emotional state of users or audiences in real time using sensors and algorithms.

[0753] This system includes terminals for user information input, a server that generates and controls progress scripts, and a mechanism for sentiment analysis. Users first input basic event information and participant characteristics via the terminal. The terminal provides an intuitive interface, making it easy for users to input information.

[0754] Information sent from the terminal is received by the server. The server uses a generative AI model to generate a natural language script based on the received information. The generated script is then translated into the required language using a multilingual translation tool. Tone and cultural appropriateness are also adjusted during this process.

[0755] For example, a possible prompt to input into a generative AI model is: "Generate a script for a 30-minute technical seminar. The target audience is young people, and it should include interactive elements."

[0756] The translated script is returned to the terminal, where the user can review and edit it as needed. Based on the user's edited script, the server adjusts the virtual presenter's voice generation and visual presentation, and prepares for the event. During this process, the server uses an emotion analysis engine to analyze the emotions of the user and audience in real time as the event progresses. Based on the emotion analysis results, the virtual presenter's expressions and tone are dynamically changed.

[0757] In this way, the system can provide users and audiences with a customized, interactive event experience. The system can also flexibly respond to unexpected situations during the event and make dynamic adjustments to ensure a realistic and engaging presentation.

[0758] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0759] Step 1:

[0760] Users use their devices to enter detailed information about the event. This information includes the date and time of the event, the theme, and the profile of the target audience. The entered information is organized as digital data on the device and sent to the server.

[0761] Step 2:

[0762] The server begins data processing based on the information received from the terminal. Specifically, it uses a generative AI model to generate a script for event progression. This model analyzes the input data based on the provided prompts and outputs natural language text. The generated script is then formatted as digital text data.

[0763] Step 3:

[0764] The server sends the generated script to a multilingual translation system. This system translates the script into the target language, evaluates and adjusts for cultural appropriateness. The translated script is then sent back to the terminal. During translation, language data conversion processing is performed, and the output is a natural expression in the target language.

[0765] Step 4:

[0766] The user reviews and edits the translated script on their device. If the user changes terminology in the script or adds new data, the device formats the edited data for resending to the server. This step involves text editing of the data.

[0767] Step 5:

[0768] The server prepares the virtual presenter based on the edited script. Specifically, it uses a speech generation engine to convert the script into audio data and a visual representation system to generate a visual presentation. In parallel, it uses an emotion analysis engine to analyze the emotions of users and the audience and generates data for real-time adjustments.

[0769] Step 6:

[0770] The terminal controls the event host in real time, receiving updated data from the server and dynamically adjusting the presentation. The system combines audio and visuals to create a realistic and interactive experience for the audience. Processing is performed to change the host's tone and gestures in response to the event situation and user feedback.

[0771] (Application Example 2)

[0772] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0773] In modern consumer activity, users require information about products and services that is tailored to their emotions and interests at that moment. However, conventional systems often lack sufficient real-time sentiment analysis and dynamic adjustment of information delivery, limiting the user experience. Furthermore, efficiently and culturally appropriate information delivery to global consumers who require multilingual support remains a challenge.

[0774] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0775] In this invention, the server includes means for providing an intuitive data input interface for users to input basic event information; means for generating an event progress script using a language processing device based on the received information; a translation device for translating the generated script into multiple languages; means for presenting the translated script to the user and providing a user-editable interface; means for controlling a virtual presenter by integrating speech synthesis and visual representation based on the edited script; an emotion analysis device for analyzing customer emotions and dynamically adjusting the information provided; and means for reflecting real-time information updates during the event. This enables personalized information provision in response to user emotions, thereby improving the user experience.

[0776] A "user" is someone who seeks information about products and services during their consumption activities.

[0777] A "data entry interface" is a means for users to intuitively input event and product information.

[0778] A "language processing device" is a device that analyzes and processes natural language and generates event progress scripts in various languages ​​based on the underlying information.

[0779] A "translation device" is a device used to translate a generated script into multiple languages.

[0780] An "interface" is a means by which the user can view and edit the provided translation script.

[0781] "Speech synthesis" is a technology that generates speech based on a processed script.

[0782] "Visual representation" refers to a means of visually representing a virtual narrator in conjunction with speech synthesis.

[0783] A "virtual facilitator" is a virtual display character that plays the role of guiding an event in a digital environment.

[0784] An "emotion analysis device" is a device that analyzes a user's emotions and dynamically adjusts the way information is provided based on that data.

[0785] "Personalization" refers to providing customized information tailored to the individual user's interests and emotions.

[0786] "Real-time" refers to something that happens immediately in the present moment.

[0787] To implement this invention, a server plays a central role. This server operates in conjunction with various software, as described below. First, the user inputs basic event information from a terminal through a data input interface. The terminal used here could be a general-purpose computer or smart device. Next, the server uses a language processing unit to perform natural language processing based on the input information and generates an event progress script. In this process, the server utilizes an advanced generative AI model.

[0788] The generated scripts are translated into multiple languages ​​using a translation tool. Services such as the Google Cloud Translation API can be used for this translation. The translated scripts are presented to the user and can be edited via a dedicated interface. This editing process also takes place on the user's device.

[0789] Once editing is complete, the server combines speech synthesis and visual representation technologies to generate a virtual presenter. Software such as Amazon Polly is used for speech synthesis, and the virtual presenter provides information to the user or audience. Furthermore, the server uses an emotion analyzer to analyze the emotions of the user or audience in real time. This analysis is then processed using machine learning libraries such as TensorFlow.

[0790] Based on the analysis results, the server dynamically adjusts how information is delivered to provide personalized information. For example, if the sentiment analyzer determines that the user seems uninterested while receiving a product description, the system will provide additional interesting information.

[0791] An example of a prompt message could be input to the generating AI model, such as, "Provide more interesting information based on the user's emotions." In this way, the present invention utilizes real-time user feedback to achieve effective information transmission.

[0792] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0793] Step 1:

[0794] The user enters basic event information into a data input interface via their device. The entered information is sent to the server as text data. The server receives this information and prepares it for natural language processing.

[0795] Step 2:

[0796] The server generates an event progress script using a language processing unit based on the received information. In this process, it utilizes a generative AI model to perform advanced natural language processing. The input is basic information provided by the user, and the output is the text data of the integrated event progress script.

[0797] Step 3:

[0798] The server uses a translation tool to translate the generated script into the required languages. The Google Cloud Translation API can be used for this purpose. The input is the text data of the generated script, and the output is the multilingual text data of the translated script.

[0799] Step 4:

[0800] The translated script is displayed on the user's terminal and becomes editable. The user reviews the information through the presented interface and makes corrections as needed. The input is the translated script, and the output is the script edited by the user.

[0801] Step 5:

[0802] Based on the edited script, the server integrates speech synthesis and visual representation to generate a virtual narrator. Amazon Polly or similar tools are used for speech synthesis, and the generated virtual narrator is displayed on the terminal. The input is the user-edited script, and the output is the virtual narrator's visual and audio data.

[0803] Step 6:

[0804] The server uses an emotion analysis device to analyze the emotions of users and audiences in real time. This analysis uses data acquired from the camera and microphone on the device, and is performed using TensorFlow or similar software. The input is the acquired emotion data, and the output is the emotion analysis result.

[0805] Step 7:

[0806] Based on the analysis results, the server dynamically adjusts how information is delivered. Specifically, the information and tone are changed according to the user's emotions. For example, the prompt "Please provide additional information to maintain the user's interest" is input into the generating AI model, and the customized information output is presented by a virtual presenter.

[0807] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0808] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0809] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0810] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0811] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0812] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0813] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0814] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0815] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0816] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0817] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0818] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0819] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0820] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0821] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0822] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0823] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0824] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0825] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0826] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0827] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0828] The following is further disclosed regarding the embodiments described above.

[0829] (Claim 1)

[0830] A means of providing an intuitive user interface for users to input basic event information,

[0831] A means for generating an event progress script using a natural language processing model based on received information,

[0832] A multilingual translation method for translating the generated script into multiple languages,

[0833] A means of presenting the translated script to the user and providing a user-editable interface,

[0834] A means for controlling a virtual presenter by integrating speech synthesis and visual representation based on an edited script,

[0835] A means of reflecting real-time information updates during the event,

[0836] A system that includes this.

[0837] (Claim 2)

[0838] The system according to claim 1, comprising tone analysis means for evaluating and adjusting the cultural relevance of the generated script.

[0839] (Claim 3)

[0840] The system according to claim 1, comprising means for providing a customization function for the virtual presenter's expression according to user requests.

[0841] "Example 1"

[0842] (Claim 1)

[0843] A means of providing an intuitive input device for users to input information,

[0844] A means for generating a progressive sentence using a language model based on received information,

[0845] A means comprising a multilingual conversion device for converting generated sentences into multiple languages,

[0846] A means for displaying the converted text to the user and providing a user-modifiable display device,

[0847] A means of controlling a virtual character by integrating sound generation and visual representation based on modified text,

[0848] A means to reflect the function of instantly updating information while it is in progress,

[0849] A system that includes this.

[0850] (Claim 2)

[0851] The system according to claim 1, comprising a tone analysis device for evaluating and adjusting the cultural consistency of the generated sentences.

[0852] (Claim 3)

[0853] The system according to claim 1, comprising means for providing an individual adjustment function for the representation of a virtual character in accordance with the user's requests.

[0854] "Application Example 1"

[0855] (Claim 1)

[0856] A means of providing an intuitive user interface for users to input basic event information,

[0857] A means for generating an event progress script using a natural language processing model based on received information,

[0858] A multilingual translation method for translating the generated script into multiple languages,

[0859] A means of presenting the translated script to the user and providing a user-editable interface,

[0860] A means for controlling a virtual presenter by integrating speech synthesis and visual representation based on an edited script,

[0861] A means of guiding visitors through a virtual store via smart glasses and presenting product information by integrating audio and visual information,

[0862] A means of reflecting real-time information updates during the event,

[0863] A means of communicating information by outputting a progress script as audio in real time based on user prompts,

[0864] A system that includes this.

[0865] (Claim 2)

[0866] The system according to claim 1, comprising tone analysis means for evaluating and adjusting the cultural relevance of the generated script, and for making adjustments to match product presentations in a virtual store.

[0867] (Claim 3)

[0868] The system according to claim 1, which provides a customization function for the virtual presenter's expression according to user requests and includes means for adjusting the visual expression in product introductions.

[0869] "Example 2 of combining an emotion engine"

[0870] (Claim 1)

[0871] A means of providing an intuitive interface using a terminal for users to input information,

[0872] A means for generating a progress script using a data processing model based on received information,

[0873] A multilingual translation tool for converting the generated script into multiple natural languages,

[0874] A means of providing an interface that presents the translated script to the user on the device and allows them to edit it,

[0875] A means for controlling a virtual facilitator by integrating speech generation and visual representation based on an edited script,

[0876] A means of updating and reflecting information in real time during the event,

[0877] A means of performing emotional analysis and dynamically adjusting the facilitator's expression according to the audience's state,

[0878] A system that includes this.

[0879] (Claim 2)

[0880] The system according to claim 1, comprising text tone analysis means for evaluating and adjusting the cultural suitability of the generated script.

[0881] (Claim 3)

[0882] The system according to claim 1, further comprising means for providing an individual adjustment function in response to user requests for the representation of the virtual facilitator.

[0883] "Application example 2 when combining with an emotional engine"

[0884] (Claim 1)

[0885] A means of providing an intuitive data entry interface for users to input basic event information,

[0886] A means for generating an event progress script using a language processing device based on received information,

[0887] A translation device for translating the generated script into multiple languages,

[0888] A means of presenting the translated script to the user and providing a user-editable interface,

[0889] A means for controlling a virtual narrator by integrating speech synthesis and visual representation based on an edited script,

[0890] An emotion analysis device that analyzes customer emotions and dynamically adjusts the information provided,

[0891] A means of reflecting real-time information updates during the event,

[0892] A system that includes this.

[0893] (Claim 2)

[0894] The system according to claim 1, comprising adjustment analysis means for evaluating and adjusting the cultural relevance of the generated script.

[0895] (Claim 3)

[0896] The system according to claim 1, comprising means for providing a customization function for the representation of a virtual facilitator in accordance with user requests. [Explanation of Symbols]

[0897] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of providing an intuitive user interface for users to input basic event information, A means for generating an event progress script using a natural language processing model based on received information, A multilingual translation tool for translating the generated script into multiple languages, A means of presenting the translated script to the user and providing a user-editable interface, A means for controlling a virtual presenter by integrating speech synthesis and visual representation based on an edited script, A means of reflecting real-time information updates during the event, A system that includes this.

2. The system according to claim 1, comprising tone analysis means for evaluating and adjusting the cultural relevance of the generated script.

3. The system according to claim 1, comprising means for providing a customization function for the virtual presenter's expression according to user requests.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A