system
The system generates and customizes virtual characters for advertising, addressing labor and consistency issues, reducing costs and risks, and enabling effective global campaigns.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-15
- Publication Date
- 2026-04-27
AI Technical Summary
Conventional advertising methods using real people face high labor costs, complexity, unpredictable risks, and difficulty in maintaining consistency due to aging and market changes, especially in global campaigns that cross language and cultural barriers.
A system for generating virtual characters that collects data, analyzes it to create a basic model, customizes the character based on company requirements, and improves performance based on feedback, enabling sustainable and effective advertising across various media platforms.
Reduces advertising costs and risks while ensuring consistent messaging across cultures and languages, allowing for responsive and efficient global campaigns.
Smart Images

Figure 2026070222000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In a company's advertising activities, when using conventional real people, there are problems such as high labor costs, complexity of contracts, unpredictable risks related to people, and difficulty in consistently using the same person due to aging and market changes. Furthermore, in global advertising campaigns, it is necessary to cross language and cultural barriers. The aim is to solve these problems and provide a sustainable and effective advertising method.
Means for Solving the Problems
[0005] This invention provides a system for generating virtual characters used in corporate advertising activities. Specifically, it first provides means for collecting data necessary for creating a virtual character, and means for analyzing that data to create a basic model of the virtual character's movements and voice. Furthermore, it provides means for customizing the virtual character based on the company's requirements, making the generated virtual character available on various media platforms. In addition, it includes means for automatically improving the performance of the virtual character based on feedback data from advertising campaigns. This achieves cost reduction and risk avoidance, enabling sustainable advertising activities that are responsive to language and culture.
[0006] A "company" is a legal entity or organization that provides specific products or services and operates for profit.
[0007] "Advertising activities" refer to means of communicating information and related processes aimed at increasing awareness of products and services and appealing to consumers.
[0008] A "virtual character" is a digital character created using computer graphics or artificial intelligence technology, intended for use in public spaces such as advertisements.
[0009] "Data" refers to recorded information elements in numerical, text, audio, image, or other forms collected through information processing technology.
[0010] "Analysis" is the process of computation or analysis that involves collecting data, processing it for a specific purpose, and extracting meaningful results.
[0011] A "basic model" is a model built based on data analysis that possesses fundamental characteristics such as the movements and voice of a virtual character.
[0012] "Customization" refers to adjusting or modifying the characteristics of a virtual character or system to suit specific needs or requirements.
[0013] A "media platform" is a foundation for distributing content through various media such as the internet, television, radio, and print.
[0014] "Feedback data" refers to data obtained from information gathering based on advertising activities and system experiences, used to evaluate performance and effectiveness.
[0015] "Improving performance" means using feedback data to make adjustments and developments that enhance the accuracy and efficiency of the system or virtual characters. [Brief explanation of the drawing]
[0016] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11]It is a sequence diagram showing the processing flow of the data processing system in Embodiment 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when combined with an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when combined with an emotion engine.
Mode for Carrying Out the Invention
[0017] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), and the like.
[0020] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0021] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0024] [First Embodiment]
[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0037] This invention provides a system for using virtual characters in a company's advertising activities and a method for supporting advertising activities based on this system. The system aims to efficiently perform advanced data analysis, generate behavioral and voice models, customize them, and deploy them in advertising activities.
[0038] First, the server collects a wide range of data necessary for generating a virtual character. This includes image data, audio data, and cultural and market background data. This data forms the basis for determining the behavior and characteristics of the virtual character.
[0039] Next, the server analyzes the collected data to create a basic model of the virtual character's movements and voice. This allows for a highly accurate reproduction of a digital character with realistic human facial expressions, movements, and vocal qualities.
[0040] Subsequently, companies customize the virtual persona to create one that suits their brand image. Through the interface, users can adjust the virtual persona's appearance, voice tone, and even its behavior in specific scenarios. This customization process allows companies to easily design a unique virtual persona that perfectly fits their market strategy.
[0041] Once the virtual character creation is complete, the device is responsible for distribution to various media platforms. The generated digital data of the virtual character can be widely applied to television commercials, online videos, social media content, and more. This allows companies to efficiently conduct global advertising campaigns.
[0042] Furthermore, after running an advertising campaign, users can send feedback data to the server. This data includes viewership ratings, engagement, and social media reactions. Based on this feedback, the server can continuously improve the performance of the virtual character and optimize it in response to new market trends.
[0043] One example is when an international beverage company launches a new product and uses a virtual character that matches its youthful image. This virtual character can deliver speeches tailored to the local language and culture, effectively conveying a consistent brand message across various countries. This allows the company to reduce advertising campaign costs while conducting effective promotions.
[0044] The following describes the processing flow.
[0045] Step 1:
[0046] The server collects the data necessary for generating virtual characters from the internet and internal databases. Specifically, it gathers image data, audio clips, text information, and data on cultural background to build the foundation for the reliability and expressiveness of the virtual characters.
[0047] Step 2:
[0048] The server analyzes the collected data and creates a basic model of the virtual character's movements and voice. Machine learning algorithms are used to extract patterns from the collected images and audio, and based on these, a model is designed to reproduce natural movements and speech patterns.
[0049] Step 3:
[0050] Users customize virtual characters through the provided interface. They can select appearances, personalities, and voices that align with the company's brand image, and set the message the virtual character should convey.
[0051] Step 4:
[0052] The server generates the final virtual character based on the customization information specified by the user. The generation process utilizes advanced 3D rendering technology to create the customized model in real time and save it as digital data.
[0053] Step 5:
[0054] The device distributes the generated virtual character to various media platforms such as television, the internet, and social media. Because the distribution is done in a format optimized for each media platform, high-quality visual presentation is possible.
[0055] Step 6:
[0056] Users send feedback data to the server to check the effectiveness of their advertising campaigns. This feedback data includes viewership ratings, viewer engagement, and comment content, which is used to measure the campaign's effectiveness.
[0057] Step 7:
[0058] The server analyzes the received feedback data and makes adjustments to improve the performance of the virtual character. The results of this analysis are reflected in the creation of future virtual characters and advertising campaigns, forming the basis for achieving greater effectiveness.
[0059] (Example 1)
[0060] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0061] In modern advertising, there is a growing need for virtual characters that can flexibly adapt to target markets and cultures. However, traditional systems have struggled to efficiently and effectively generate, customize, distribute, and improve these virtual characters. Therefore, there is a need for methods that maximize advertising effectiveness while keeping costs down.
[0062] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0063] In this invention, the server includes means for collecting a wide range of information necessary for generating a virtual character, means for analyzing the collected information and creating the basic structure of the virtual character's movements and voice, and means for adjusting the virtual character to match the company's brand image. This makes it possible to efficiently generate a virtual character optimized for the target market and utilize it in various media.
[0064] A "virtual character" is a computer-generated character created using digital technology that can behave like a real person.
[0065] "Information" refers to a dataset containing image data, audio data, and cultural and market background data necessary for generating a virtual character.
[0066] "Device" refers to a combination of hardware and software used to collect, analyze, and process information according to a specific purpose.
[0067] "Analysis" is a data processing process performed to derive the basic structure of a virtual character's movements and voice from the collected data.
[0068] "Basic structure" refers to the fundamental frameworks and algorithms necessary to represent the actions and voices of virtual characters.
[0069] "Brand image" refers to the characteristic of a company that uses fictional characters to express its philosophy and values and appeal to consumers.
[0070] "Adjustment" is the process of modifying a virtual character's appearance, voice, and behavior to meet specific requirements.
[0071] This invention is a system aimed at generating and applying virtual characters in corporate advertising activities. The system aims to maximize advertising effectiveness by analyzing various data, creating customizable virtual characters, and distributing them across multiple media platforms.
[0072] The server collects a wide range of information necessary for generating virtual characters. This process involves using web scraping tools and database management systems to gather images, audio, and data related to cultural background. Specific software examples include tools like Beautiful Soup and Scrapy, which are used for data collection.
[0073] Next, the server analyzes the acquired information and uses machine learning algorithms to create the basic structure of the virtual character's movements and voice. Here, machine learning libraries such as TENSORFLOW® and PyTorch are used to train the model. This generates a character that behaves like a real human.
[0074] Based on the generated foundational structure, users can customize the virtual character to match the company's brand image through a dedicated interface. The UI is designed to be user-friendly, allowing users to adjust appearance, voice tone, and behavior in specific scenarios using sliders and menus.
[0075] Subsequently, the device processes the customized virtual character's digital data and distributes it in the appropriate format to a wide variety of media, including television commercials, online video platforms, and social media. This process includes converting the data using video encoding tools such as FFmpeg and distributing it to media platforms using APIs.
[0076] Finally, users collect feedback on advertising activities, send it to the server, and use it to improve the performance of the virtual character. Data analysis tools such as Google Analytics are used to analyze the feedback data, allowing the virtual character to continuously improve in response to market trends.
[0077] A concrete example would be a company using this system to generate a virtual character that matches the image of a new product when promoting it, and then launching an advertising campaign in the international market. In this case, an example of a prompt might be, "Please generate a youthful and approachable character to support the new product launch campaign."
[0078] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0079] Step 1:
[0080] The server collects a wide range of information necessary for generating virtual characters. As input, it obtains relevant image data, audio data, and cultural / market background data from the internet and internal databases. This data is collected using web scraping tools (e.g., Beautiful Soup, Scrapy). As output, it provides the raw dataset necessary for creating the base model of the virtual character.
[0081] Step 2:
[0082] The server analyzes the collected data and generates the basic structure of the virtual character's movements and voice. The input is the dataset obtained in Step 1. For analysis, machine learning libraries such as TensorFlow and PyTorch are used to train image recognition models and speech synthesis models. As output, basic structure data of a virtual character with realistic movements and voice is generated.
[0083] Step 3:
[0084] The user customizes the virtual character to match the company's brand image using a dedicated interface. The input is the basic structure data generated in step 2. The user manipulates sliders and menus to adjust the virtual character's appearance, voice tone, and specific behaviors. The output is the configuration data of the customized virtual character.
[0085] Step 4:
[0086] The terminal converts the customized virtual person into a digital format suitable for various media formats and distributes it. The input is the configuration data created in step 3. Using a video encoding tool such as FFmpeg, the digital data of the virtual person is converted into a format suitable for television commercials, online videos, and social media content. The output is content in a format suitable for each media platform.
[0087] Step 5:
[0088] Users collect feedback data on their advertising activities and send it to the server. The input is the results of advertising activities, such as viewership and engagement. Data analysis tools such as Google Analytics are used for analysis, and performance evaluations and improvement suggestions for the virtual character are generated. The output is feedback data provided for future updates and improvements to the virtual character.
[0089] (Application Example 1)
[0090] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0091] In modern advertising, there is a growing demand for more personalized messaging and promotions tailored to users' interests and preferences. Traditional methods often result in static, generalized advertisements, making it difficult to efficiently deliver compelling messages to individual users. Furthermore, the inability to quickly utilize feedback from advertising campaigns leads to limited advertising effectiveness.
[0092] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0093] In this invention, the server includes means for collecting information to generate a virtual character to be used in a medium; means for analyzing the collected information to create a basic model of the virtual character's movements and voice; means for adjusting the virtual character based on the requirements of the medium; means for generating the final virtual character based on the adjusted settings; means for distributing the generated virtual character to multiple information distribution platforms; means for improving the performance of the virtual character based on feedback data; and means for customizing the virtual character based on user interests and effectively conveying messages through videos and interactive simulations. This makes it possible to provide advertisements optimized for each user and maximize the effectiveness of advertising activities.
[0094] A "medium" is a method or technical means for transmitting a message over a wide area.
[0095] A "virtual character" is a character created using digital technology that imitates human characteristics and behaviors.
[0096] "Information" refers to the data necessary for creating and customizing virtual characters, and is handled at each stage of collection, analysis, and application.
[0097] A "server" is a computer system that stores, analyzes, and processes data, and supports the process of creating virtual characters.
[0098] A "distribution platform" is an infrastructure for distributing content and providing information, and it plays a role in conveying messages from virtual characters to users.
[0099] "Feedback data" refers to evaluation information regarding the performance of advertising campaigns and virtual characters, and serves as a basis for improvement and adjustments.
[0100] "Customization" is the act of adjusting the appearance and behavior of a virtual character to suit specific requirements and preferences, and is an important process for meeting user expectations.
[0101] "Simulation" is a technology that recreates a virtual situation on a computer and predicts specific results or effects, providing an interactive experience for the user.
[0102] A "message" is the information and brand image delivered to users through advertising activities, and is the central content of the communication.
[0103] The system for implementing this invention streamlines the process of generating virtual characters and utilizing them in advertising activities. The server first collects information necessary to generate virtual characters to be used in the media, such as image data, audio data, and market background information. Based on this collected information, it builds a basic model to mimic the actions and voice of the virtual character. In this process, it uses TensorFlow, a machine learning platform, to perform data processing and calculations to reproduce the natural behavior and voice characteristics of the character.
[0104] The device provides the ability to customize virtual characters through user interaction. Users can intuitively adjust the appearance and voice tone of the virtual character through the interface. Using Unity, it is possible to render the virtual character's movements in real time and provide visual feedback. Specifically, by using the app and adjusting the virtual character to suit a particular scenario, a character optimized for on-the-spot communication can be generated.
[0105] The digital data of virtual characters generated through the distribution platform is widely deployed in online advertising, social media content, and other areas. This process makes it possible to efficiently implement global advertising campaigns.
[0106] Furthermore, users send feedback to the server regarding the performance of their advertising campaigns. This feedback includes viewership ratings, engagement, and social media reactions, and the analysis results are used to improve the performance of virtual characters and optimize them in response to market trends.
[0107] As a concrete example, a prompt for generating a virtual character who introduces casual fashion for women in their 20s could be: "Please generate a virtual character who introduces popular casual fashion items for women in their 20s. The character should have a bright expression, greet in a friendly tone, and clearly explain the features of the products."
[0108] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0109] Step 1:
[0110] The server collects the information necessary to generate virtual characters used in the media. This includes image data, audio data, and cultural and market background information from multiple data sources. The input is this diverse set of data, and the output is a basic dataset for virtual character generation. Specifically, it automates data acquisition through web scraping and APIs.
[0111] Step 2:
[0112] The server analyzes the collected information and creates a basic model of the virtual character's movements and speech. It processes the data using TensorFlow and builds a machine learning model. The input is the basic dataset, and the output is the virtual character's movements and speech model. Specifically, it performs data cleaning, feature selection, and model training.
[0113] Step 3:
[0114] Users customize their virtual character through the device's interface. This includes adjusting appearance, voice tone, and behavior. Input is configuration information from the base model and user interface, and output is the customized virtual character model. Specifically, intuitive configuration changes are possible via the UI.
[0115] Step 4:
[0116] The device uses Unity to render virtual characters in real time and provide visual feedback. The input is a customized virtual character model, and the output is an interactive character displayed to the user. Specifically, it handles the rendering of 3D models and user interaction.
[0117] Step 5:
[0118] The server sends the generated virtual character's digital data to the distribution platform for distribution as online advertisements and social media content. The input is the completed virtual character model data, and the output is its display on the distribution platform. Specifically, it configures content distribution settings through the APIs of each platform.
[0119] Step 6:
[0120] Users send feedback to the server regarding the performance of their advertising campaigns. This feedback includes data such as viewership and engagement; the input is the campaign results, and the output is an analytical report for improvement. Specifically, data analysis software is used to evaluate the results and propose improvements.
[0121] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0122] This invention relates to a virtual character generation system that incorporates an emotion engine, with the aim of achieving more personalized effects in corporate advertising activities. This system has the ability to recognize the user's emotions in real time and automatically adjust the behavior and voice expression of the virtual character accordingly.
[0123] First, as an embodiment, the server collects and analyzes video and audio data for emotion recognition. Specifically, it measures the user's facial expressions, voice tone, and tempo, and analyzes this data using a machine learning algorithm to identify the user's emotional state. The collected data is used to build an emotion model, which is then reflected in the behavior of the virtual character.
[0124] Next, the server uses an emotion engine to make adjustments to the virtual character's base model that correspond to the user's emotions. These adjustments allow the virtual character to display changes in facial expressions and voice quality in accordance with the user's emotions, enabling more realistic communication. For example, when the user expresses happiness, the virtual character is adjusted to speak with a brighter expression and a more cheerful tone.
[0125] Furthermore, users can further customize the virtual character to match the characteristics of the company's brand. They can use the interface to set the appearance and personality of the virtual character, providing a personalized user experience tailored to the company's marketing objectives.
[0126] The generated virtual characters are distributed to various media platforms via the device. When distributed through television, online advertising, social media, etc., the aforementioned emotional response is utilized in interactions with actual users. This makes it possible to further enhance the effectiveness of advertising.
[0127] One possible use case is an international cosmetics company using this system to promote a new product. In this scenario, when a target user interacts through the camera, their emotions, such as interest or surprise, are identified, and a virtual character responds by providing product information and messages accordingly. This provides a more personalized advertising experience, increasing the user's interest in the product and their willingness to purchase it.
[0128] The following describes the processing flow.
[0129] Step 1:
[0130] The server collects the data necessary to recognize the user's emotions. Specifically, it acquires the user's facial expression and voice data through cameras and microphones, and inputs this data into the emotion analysis system.
[0131] Step 2:
[0132] The server uses machine learning algorithms to analyze the user's emotions based on acquired facial and voice data. Here, emotions are identified using elements such as smiles, the degree of frowning, and the tone and rhythm of the voice, and the emotional state is recognized in real time.
[0133] Step 3:
[0134] The server dynamically adjusts the virtual character's behavior and voice expression based on the results of emotion analysis. For example, if the user indicates feelings of joy, the server instructs the virtual character to respond with a cheerful expression and a friendly tone.
[0135] Step 4:
[0136] Users customize their virtual avatars through the interface. They can set the appearance and style of their virtual avatars to match specific advertising campaigns, providing an experience tailored to their individual preferences.
[0137] Step 5:
[0138] The server generates a final model of the virtual character based on customization and emotion recognition, and stores it as digital data suitable for use. This enables consistent representation of the virtual character across various media platforms.
[0139] Step 6:
[0140] The device distributes the generated virtual character to multiple media platforms, including television, online advertising, and social media. The virtual character is used as part of digital content, enabling effective communication that reflects the user's emotional responses.
[0141] Step 7:
[0142] Users provide feedback on the performance of advertising campaigns and user response data, and the server analyzes this data to further improve the performance of the virtual character and emotion engine. Through this process, the ability to more accurately target the audience in the next campaign is enhanced.
[0143] (Example 2)
[0144] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0145] Traditional advertising methods have struggled to deliver personalized ads that accurately capture the emotions of target users, resulting in limited advertising effectiveness. In particular, providing a consistent customer experience across diverse media platforms is difficult, and the inability to utilize real-time feedback that reflects user experience is a significant problem.
[0146] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0147] In this invention, the server includes means for analyzing video and audio information acquired from the user to estimate the user's emotional state, means for automatically adjusting the actions and voice quality of the virtual character based on the estimated emotional state, and means for generating a user's emotional model and reflecting it in the virtual character's expression. This makes it possible to provide a dynamic and personalized advertising experience tailored to the user.
[0148] "Video and audio information acquired from the user" refers to visual and auditory data collected through the user's video camera and microphone, which forms the basis for understanding the user's emotional state.
[0149] "Means for estimating emotional states" refers to algorithms or technologies that analyze collected video and audio information to identify the emotions expressed by the user.
[0150] "Means for automatically adjusting the actions and voice quality of a virtual character" refers to technology that changes the facial expressions and voice quality of a virtual character in real time according to the user's estimated emotional state.
[0151] "A means of generating a user emotion model and reflecting it in the representation of a virtual character" refers to a technology that creates a database of user emotion tendencies and then uses that information to create a model that reflects the actions and speech of a virtual character.
[0152] A "personalized advertising experience" is a method of providing users with an individually optimized experience by presenting advertisements whose content dynamically changes according to each user's preferences and emotions.
[0153] This invention is a system that analyzes a user's emotional state and provides a personalized advertising experience through a virtual character based on that analysis. This system processes the user's video and audio information in real time to detect emotions and reflect them in the virtual character's actions and voice quality.
[0154] First, the server acquires video and audio information from the user's device. For video analysis, visual data processing software such as "OpenCV" is used for face recognition and facial expression analysis. For audio analysis, audio analysis tools such as "Praat" are used to analyze the tone and tempo of the sound. From this information, a machine learning algorithm (e.g., TensorFlow) is applied to estimate the user's emotional state.
[0155] Next, based on the estimated emotions, the server uses an emotion engine to adjust the base model of the virtual character. This allows the virtual character's facial expressions and voice quality to change in real time according to the user's emotions. As a result, the user can feel as if they are interacting with a real person.
[0156] Users can customize the appearance and personality of their virtual characters through the provided interface. This allows for the creation of virtual characters with characteristics tailored to the user's needs and the company's advertising goals.
[0157] Ultimately, the refined virtual persona is distributed via the device to various media outlets such as television, internet advertising, and social media. This makes it possible to provide a customized advertising experience for each user, leading to the expectation of higher advertising effectiveness.
[0158] For example, if a user smiles, the server quickly picks up on that emotion and adjusts the virtual character to give a positive response such as, "That's good news!" This improves the quality of the ad interaction.
[0159] Examples of prompts include, "How should the virtual character's expression change when the user's facial expression changes?" and "How should the virtual character speak to a user who is feeling anxious?" By using these prompts, the AI model can learn to achieve more natural dialogue.
[0160] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0161] Step 1:
[0162] The server retrieves video and audio data from the user.
[0163] The input consists of video and audio information acquired in real time from the user's camera and microphone. The server receives this data as a stream and prepares it for analysis. Specifically, the data is sent to the server via an API.
[0164] Step 2:
[0165] The server analyzes video and audio data to estimate the user's emotional state.
[0166] The input consists of video and audio data acquired in Step 1. The server uses machine learning algorithms (e.g., TensorFlow) to perform analysis, using "OpenCV" for facial expression analysis and "Praat" for audio analysis. The output is an emotional state profile that represents the user's emotions. Specifically, the server extracts facial features and applies algorithms to measure the tone and tempo of the audio.
[0167] Step 3:
[0168] The server handles the adjustments of the virtual characters.
[0169] As input, the system receives an emotional state profile from Step 2. The server uses a generative AI model and leverages an emotion engine to refine the base model of the virtual character. The output is the refined virtual character data, which includes changes in facial expressions and voice. In terms of specific actions, the server modifies the virtual character's movements and voice quality in real time according to the emotional profile.
[0170] Step 4:
[0171] Users customize their virtual characters.
[0172] The input consists of customization information provided by the user through the operation screen. Users can set the appearance, personality, voice quality, etc., of a virtual character, specifying characteristics that match the company's advertising goals. The output is the user-customized settings information of the virtual character. Specifically, changes specified by the user are reflected on the server through interaction via the GUI.
[0173] Step 5:
[0174] The device distributes virtual characters to media platforms.
[0175] The input is data for a virtual person, which is adjusted in step 3 and customized in step 4. The device delivers this data to users through television, social media, online advertising, etc. The output is the user's visual and auditory experience. Specifically, the device is integrated with existing streaming services and advertising delivery systems to deliver the virtual person in real time.
[0176] (Application Example 2)
[0177] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0178] In today's world, simply presenting products is often insufficient for corporate advertising to effectively appeal to consumers. In particular, recognizing the emotional state of individual consumers and presenting the most appropriate advertisements accordingly is an effective approach. However, conventional systems have been inadequate in real-time emotion recognition and the dynamic adjustment of advertisements based on that recognition. Therefore, there is a need to build systems that provide a more personalized advertising experience based on consumer emotions.
[0179] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0180] In this invention, the server includes means for collecting data to recognize a person's emotional state in real time and dynamically adjust the advertising content according to that emotional state; means for analyzing the collected data to create a basic model of the virtual person's actions and voice; and means for adjusting the virtual person's actions and voice based on the emotion analysis data. This makes it possible to dynamically adjust the advertising content based on the emotions of each individual consumer and provide effective advertising in real time.
[0181] "A person's emotional state" refers to an individual's inner feelings and mood, as perceived from their facial expressions, voice, and other cues.
[0182] "Real-time recognition" refers to a process where information is processed as soon as it is generated, and results are obtained immediately.
[0183] "Dynamically adjusting ad content" means instantly changing the information presented as an ad according to the user's situation and emotions.
[0184] "Means of data collection" refers to any technical device or method for acquiring audio, image, or other sensory data.
[0185] A "foundation model" refers to a basic system or framework built to provide a foundation for a specific application.
[0186] "Emotional analysis data" refers to the results of data analysis used to identify a user's emotional state.
[0187] "Means for adjusting the behavior and voice of a virtual character" refers to a mechanism for changing the movements and speech patterns of a generated virtual character.
[0188] "Feedback data" refers to data collected from users' reactions and opinions to the information provided by the system.
[0189] A "media platform" is an online or offline technological infrastructure for distributing information and content.
[0190] A server plays a central role in implementing this invention. The server first collects the user's facial expressions and voice data through smart glasses or other devices. This data is analyzed in real time using emotion recognition software such as Google Cloud Vision or Amazon Rekognition to identify the user's emotional state. The analysis results are instantly reflected in the base model of the virtual person associated with each user.
[0191] On the server, the virtual character's actions and voice are adjusted based on analyzed emotion data. Using a generative AI model, the virtual character's behavior is designed to dynamically adjust the ad content according to the user's emotions. This virtual character is projected in real time onto smart glasses displays and other device screens, displaying ads in an interactive manner.
[0192] Furthermore, users can customize their virtual avatars through the interface, setting their appearance, voice, and behavior according to the company's marketing strategy. Feedback data is also collected and stored on the server for improvement.
[0193] For example, if a user wearing smart glasses is detected as being in a relaxed state, advertisements for calming relaxation products will be displayed with a bright and gentle voice. In this way, personalized advertising tailored to each user's emotional state is possible.
[0194] An example of a prompt for a generative AI model is, "Please think of the best advertisement to display when the user is relaxed."
[0195] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0196] Step 1:
[0197] The server collects video and audio data in real time from the user's smart glasses via the camera and microphone. This input data includes the user's facial expressions, voice tone, and tempo. This data forms the basis for analyzing the user's emotions.
[0198] Step 2:
[0199] The server inputs the collected video and audio data into emotion recognition software such as Google Cloud Vision or Amazon Rekognition to analyze the user's emotional state. This process uses machine learning algorithms to output emotion labels such as "happiness" and "stress" from the input data. This is then incorporated into the base model and used to adjust the virtual character.
[0200] Step 3:
[0201] The server uses a generative AI model to adjust the behavior and voice of the virtual character based on the analyzed emotional state. For example, if the user is relaxed, it generates a virtual character with a cheerful and calm voice and facial expression. This process also influences the selection of advertising content, leading to the output of optimal advertisements.
[0202] Step 4:
[0203] The device displays a customized virtual person and advertisement content on the smart glasses' display. In this step, the generated virtual person interacts with the user, creating a personalized advertising experience. Based on the output information from the server, the smart glasses play high-resolution video and audio.
[0204] Step 5:
[0205] Through the displayed advertising interface, users input settings for a virtual character and provide feedback on the ad content on their device. This feedback data is sent to the server and used to improve personalization in the future, contributing to overall system performance improvements.
[0206] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0207] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0208] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0209] [Second Embodiment]
[0210] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0211] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0212] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0213] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0214] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0215] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0216] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0217] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0218] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0219] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0220] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0221] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0222] This invention provides a system for using virtual characters in a company's advertising activities and a method for supporting advertising activities based on this system. The system aims to efficiently perform advanced data analysis, generate behavioral and voice models, customize them, and deploy them in advertising activities.
[0223] First, the server collects a wide range of data necessary for generating a virtual character. This includes image data, audio data, and cultural and market background data. This data forms the basis for determining the behavior and characteristics of the virtual character.
[0224] Next, the server analyzes the collected data to create a basic model of the virtual character's movements and voice. This allows for a highly accurate reproduction of a digital character with realistic human facial expressions, movements, and vocal qualities.
[0225] Subsequently, companies customize the virtual persona to create one that suits their brand image. Through the interface, users can adjust the virtual persona's appearance, voice tone, and even its behavior in specific scenarios. This customization process allows companies to easily design a unique virtual persona that perfectly fits their market strategy.
[0226] Once the virtual character creation is complete, the device is responsible for distribution to various media platforms. The generated digital data of the virtual character can be widely applied to television commercials, online videos, social media content, and more. This allows companies to efficiently conduct global advertising campaigns.
[0227] Furthermore, after running an advertising campaign, users can send feedback data to the server. This data includes viewership ratings, engagement, and social media reactions. Based on this feedback, the server can continuously improve the performance of the virtual character and optimize it in response to new market trends.
[0228] One example is when an international beverage company launches a new product and uses a virtual character that matches its youthful image. This virtual character can deliver speeches tailored to the local language and culture, effectively conveying a consistent brand message across various countries. This allows the company to reduce advertising campaign costs while conducting effective promotions.
[0229] The following describes the processing flow.
[0230] Step 1:
[0231] The server collects the data necessary for generating virtual characters from the internet and internal databases. Specifically, it gathers image data, audio clips, text information, and data on cultural background to build the foundation for the reliability and expressiveness of the virtual characters.
[0232] Step 2:
[0233] The server analyzes the collected data and creates a basic model of the virtual character's movements and voice. Machine learning algorithms are used to extract patterns from the collected images and audio, and based on these, a model is designed to reproduce natural movements and speech patterns.
[0234] Step 3:
[0235] Users customize virtual characters through the provided interface. They can select appearances, personalities, and voices that align with the company's brand image, and set the message the virtual character should convey.
[0236] Step 4:
[0237] The server generates the final virtual character based on the customization information specified by the user. The generation process utilizes advanced 3D rendering technology to create the customized model in real time and save it as digital data.
[0238] Step 5:
[0239] The device distributes the generated virtual character to various media platforms such as television, the internet, and social media. Because the distribution is done in a format optimized for each media platform, high-quality visual presentation is possible.
[0240] Step 6:
[0241] Users send feedback data to the server to check the effectiveness of their advertising campaigns. This feedback data includes viewership ratings, viewer engagement, and comment content, which is used to measure the campaign's effectiveness.
[0242] Step 7:
[0243] The server analyzes the received feedback data and makes adjustments to improve the performance of the virtual character. The results of this analysis are reflected in the creation of future virtual characters and advertising campaigns, forming the basis for achieving greater effectiveness.
[0244] (Example 1)
[0245] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0246] In modern advertising, there is a growing need for virtual characters that can flexibly adapt to target markets and cultures. However, traditional systems have struggled to efficiently and effectively generate, customize, distribute, and improve these virtual characters. Therefore, there is a need for methods that maximize advertising effectiveness while keeping costs down.
[0247] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0248] In this invention, the server includes means for collecting a wide range of information necessary for generating a virtual character, means for analyzing the collected information and creating the basic structure of the virtual character's movements and voice, and means for adjusting the virtual character to match the company's brand image. This makes it possible to efficiently generate a virtual character optimized for the target market and utilize it in various media.
[0249] A "virtual character" is a computer-generated character created using digital technology that can behave like a real person.
[0250] "Information" refers to a dataset containing image data, audio data, and cultural and market background data necessary for generating a virtual character.
[0251] "Device" refers to a combination of hardware and software used to collect, analyze, and process information according to a specific purpose.
[0252] "Analysis" is a data processing process performed to derive the basic structure of a virtual character's movements and voice from the collected data.
[0253] "Basic structure" refers to the fundamental frameworks and algorithms necessary to represent the actions and voices of virtual characters.
[0254] "Brand image" refers to the characteristic of a company that uses fictional characters to express its philosophy and values and appeal to consumers.
[0255] "Adjustment" is the process of modifying a virtual character's appearance, voice, and behavior to meet specific requirements.
[0256] This invention is a system aimed at generating and applying virtual characters in corporate advertising activities. The system aims to maximize advertising effectiveness by analyzing various data, creating customizable virtual characters, and distributing them across multiple media platforms.
[0257] The server collects a wide range of information necessary for generating virtual characters. This process involves using web scraping tools and database management systems to gather images, audio, and data related to cultural background. Specific software examples include tools like Beautiful Soup and Scrapy, which are used for data collection.
[0258] Next, the server analyzes the acquired information and uses machine learning algorithms to create the basic structure of the virtual character's movements and voice. Here, machine learning libraries such as TensorFlow and PyTorch are used to train the model. This generates a character that behaves like a real human.
[0259] Based on the generated foundational structure, users can customize the virtual character to match the company's brand image through a dedicated interface. The UI is designed to be user-friendly, allowing users to adjust appearance, voice tone, and behavior in specific scenarios using sliders and menus.
[0260] Subsequently, the device processes the customized virtual character's digital data and distributes it in the appropriate format to a wide variety of media, including television commercials, online video platforms, and social media. This process includes converting the data using video encoding tools such as FFmpeg and distributing it to media platforms using APIs.
[0261] Finally, users collect feedback on advertising activities, send it to the server, and use it to improve the performance of the virtual character. Data analysis tools such as Google Analytics are used to analyze the feedback data, which allows the virtual character to continuously improve in response to market trends.
[0262] A concrete example would be a company using this system to generate a virtual character that matches the image of a new product when promoting it, and then launching an advertising campaign in the international market. In this case, an example of a prompt might be, "Please generate a youthful and approachable character to support the new product launch campaign."
[0263] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0264] Step 1:
[0265] The server collects a wide range of information necessary for generating virtual characters. As input, it obtains relevant image data, audio data, and cultural / market background data from the internet and internal databases. This data is collected using web scraping tools (e.g., Beautiful Soup, Scrapy). As output, it provides the raw dataset necessary for creating the base model of the virtual character.
[0266] Step 2:
[0267] The server analyzes the collected data and generates the basic structure of the virtual character's movements and voice. The input is the dataset obtained in Step 1. For analysis, machine learning libraries such as TensorFlow and PyTorch are used to train image recognition models and speech synthesis models. As output, basic structure data of a virtual character with realistic movements and voice is generated.
[0268] Step 3:
[0269] The user customizes the virtual character to match the company's brand image using a dedicated interface. The input is the basic structure data generated in step 2. The user manipulates sliders and menus to adjust the virtual character's appearance, voice tone, and specific behaviors. The output is the configuration data of the customized virtual character.
[0270] Step 4:
[0271] The terminal converts the customized virtual person into a digital format suitable for various media formats and distributes it. The input is the configuration data created in step 3. Using a video encoding tool such as FFmpeg, the digital data of the virtual person is converted into a format suitable for television commercials, online videos, and social media content. The output is content in a format suitable for each media platform.
[0272] Step 5:
[0273] Users collect feedback data on their advertising activities and send it to the server. The input is the results of advertising activities, such as viewership and engagement. Data analysis tools such as Google Analytics are used for analysis, and performance evaluations and improvement suggestions for the virtual character are generated. The output is feedback data provided for future updates and improvements to the virtual character.
[0274] (Application Example 1)
[0275] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0276] In modern advertising, there is a growing demand for more personalized messaging and promotions tailored to users' interests and preferences. Traditional methods often result in static, generalized advertisements, making it difficult to efficiently deliver compelling messages to individual users. Furthermore, the inability to quickly utilize feedback from advertising campaigns leads to limited advertising effectiveness.
[0277] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0278] In this invention, the server includes means for collecting information to generate virtual characters used in media, means for analyzing the collected information to create a basic model of the actions and voices of the virtual characters, means for adjusting the virtual characters based on the requirements of the media, means for generating the final virtual characters based on the adjusted settings, means for distributing the generated virtual characters to multiple information distribution platforms, means for improving the performance of the virtual characters based on feedback data, and means for customizing the virtual characters based on the interests of the users and effectively transmitting messages through videos and interactive simulations. As a result, it becomes possible to provide optimized advertisements for each user, and the effectiveness of advertising activities can be maximized.
[0279] "Media" refers to a method or technical means for transmitting messages over a wide range.
[0280] "Virtual character" refers to a character that imitates human characteristics and behaviors, generated using digital technology.
[0281] "Information" refers to the data necessary for the generation and customization of virtual characters, and is handled at each stage of collection, analysis, and application.
[0282] "Server" refers to a computer system that accumulates, analyzes, and processes data, and supports the generation process of virtual characters.
[0283] "Distribution platform" refers to the infrastructure for distributing content and providing information, and plays a role in transmitting the messages of virtual characters to users.
[0284] "Feedback data" refers to evaluation information regarding advertising campaigns and the performance of virtual characters, and serves as a criterion for improvement and adjustment.
[0285] "Customization" refers to the act of adjusting the appearance and behavior of virtual characters according to specific requirements and preferences, and it is an important process to meet users' expectations.
[0286] "Simulation" is a technology that reproduces virtual situations on a computer and predicts specific results and impacts, and it provides an interactive experience for users.
[0287] "Message" refers to the information and brand image delivered to users through advertising activities, and it is the core content of communication.
[0288] The system for implementing this invention is to generate virtual characters and streamline the process of utilizing them in advertising activities. The server first collects information necessary to generate virtual characters used in the media, such as image data, audio data, and market background information. Based on this collected information, a basic model for imitating the actions and voices of virtual characters is constructed. In this process, TensorFlow, a machine learning platform, is used to perform data processing and operations for reproducing the natural behavior and voice characteristics of people.
[0289] The terminal provides a function to customize virtual characters through interaction with users. Users can intuitively adjust the appearance and voice tone of virtual characters via the interface. Using Unity, it is possible to render the actions of virtual characters in real time and provide visual feedback. As a specific example, by using an app and adjusting virtual characters according to a specific scenario, characters optimized for on-the-spot communication can be generated.
[0290] Through the distribution platform, the digital data of the generated virtual characters is widely deployed in online advertisements, SNS content, etc. Through this process, it becomes possible to efficiently implement global advertising campaigns.
[0291] Furthermore, users send feedback to the server regarding the performance of their advertising campaigns. This feedback includes viewership ratings, engagement, and social media reactions, and the analysis results are used to improve the performance of virtual characters and optimize them in response to market trends.
[0292] As a concrete example, a prompt for generating a virtual character who introduces casual fashion for women in their 20s could be: "Please generate a virtual character who introduces popular casual fashion items for women in their 20s. The character should have a bright expression, greet in a friendly tone, and clearly explain the features of the products."
[0293] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0294] Step 1:
[0295] The server collects the information necessary to generate virtual characters used in the media. This includes image data, audio data, and cultural and market background information from multiple data sources. The input is this diverse set of data, and the output is a basic dataset for virtual character generation. Specifically, it automates data acquisition through web scraping and APIs.
[0296] Step 2:
[0297] The server analyzes the collected information and creates a basic model of the virtual character's movements and speech. It processes the data using TensorFlow and builds a machine learning model. The input is the basic dataset, and the output is the virtual character's movements and speech model. Specifically, it performs data cleaning, feature selection, and model training.
[0298] Step 3:
[0299] Users customize their virtual character through the device's interface. This includes adjusting appearance, voice tone, and behavior. Input is configuration information from the base model and user interface, and output is the customized virtual character model. Specifically, intuitive configuration changes are possible via the UI.
[0300] Step 4:
[0301] The device uses Unity to render virtual characters in real time and provide visual feedback. The input is a customized virtual character model, and the output is an interactive character displayed to the user. Specifically, it handles the rendering of 3D models and user interaction.
[0302] Step 5:
[0303] The server sends the generated virtual character's digital data to the distribution platform for distribution as online advertisements and social media content. The input is the completed virtual character model data, and the output is its display on the distribution platform. Specifically, it configures content distribution settings through the APIs of each platform.
[0304] Step 6:
[0305] Users send feedback to the server regarding the performance of their advertising campaigns. This feedback includes data such as viewership and engagement; the input is the campaign results, and the output is an analytical report for improvement. Specifically, data analysis software is used to evaluate the results and propose improvements.
[0306] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0307] The present invention relates to a virtual character generation system combined with an emotion engine, and aims to achieve a more personalized effect in the advertising activities of enterprises. This system has the ability to recognize the emotions of users in real time and automatically adjust the actions and voice expressions of virtual characters accordingly.
[0308] First, as an embodiment, the server collects and analyzes video and audio data for emotion recognition. Specifically, it measures the facial expressions, voice tones and tempos of users, and analyzes these data by machine learning algorithms to identify the emotional states of users. The collected data is used to build an emotion model and reflected in the behavior of virtual characters.
[0309] Next, the server uses the emotion engine to make adjustments to the basic model of the virtual character corresponding to the emotions of the user. Through this adjustment, the virtual character shows changes in expression and voice quality according to the emotions of the user, enabling more realistic communication. For example, when the user shows a happy emotion, the virtual character is adjusted to speak with a bright expression and an energetic tone.
[0310] In addition, the user can further customize the virtual character according to the characteristics of the corporate brand. Use the interface to set the appearance and personality of the virtual character, and provide a personalized user experience that matches the marketing goals of the enterprise.
[0311] The generated virtual character is distributed to various media platforms via a terminal. When distributed through TV, online advertising, SNS, etc., the above-mentioned emotion response is utilized in the interaction with actual users. This can further enhance the advertising effect.
[0312] One possible use case is an international cosmetics company using this system to promote a new product. In this scenario, when a target user interacts through the camera, their emotions, such as interest or surprise, are identified, and a virtual character responds by providing product information and messages accordingly. This provides a more personalized advertising experience, increasing the user's interest in the product and their willingness to purchase it.
[0313] The following describes the processing flow.
[0314] Step 1:
[0315] The server collects the data necessary to recognize the user's emotions. Specifically, it acquires the user's facial expression and voice data through cameras and microphones, and inputs this data into the emotion analysis system.
[0316] Step 2:
[0317] The server uses machine learning algorithms to analyze the user's emotions based on acquired facial and voice data. Here, emotions are identified using elements such as smiles, the degree of frowning, and the tone and rhythm of the voice, and the emotional state is recognized in real time.
[0318] Step 3:
[0319] The server dynamically adjusts the virtual character's behavior and voice expression based on the results of emotion analysis. For example, if the user indicates feelings of joy, the server instructs the virtual character to respond with a cheerful expression and a friendly tone.
[0320] Step 4:
[0321] Users customize their virtual avatars through the interface. They can set the appearance and style of their virtual avatars to match specific advertising campaigns, providing an experience tailored to their individual preferences.
[0322] Step 5:
[0323] The server generates a final model of the virtual character based on customization and emotion recognition, and stores it as digital data suitable for use. This enables consistent representation of the virtual character across various media platforms.
[0324] Step 6:
[0325] The device distributes the generated virtual character to multiple media platforms, including television, online advertising, and social media. The virtual character is used as part of digital content, enabling effective communication that reflects the user's emotional responses.
[0326] Step 7:
[0327] Users provide feedback on the performance of advertising campaigns and user response data, and the server analyzes this data to further improve the performance of the virtual character and emotion engine. Through this process, the ability to more accurately target the audience in the next campaign is enhanced.
[0328] (Example 2)
[0329] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0330] Traditional advertising methods have struggled to deliver personalized ads that accurately capture the emotions of target users, resulting in limited advertising effectiveness. In particular, providing a consistent customer experience across diverse media platforms is difficult, and the inability to utilize real-time feedback that reflects user experience is a significant problem.
[0331] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0332] In this invention, the server includes means for analyzing video and audio information acquired from the user to estimate the user's emotional state, means for automatically adjusting the actions and voice quality of the virtual character based on the estimated emotional state, and means for generating a user's emotional model and reflecting it in the virtual character's expression. This makes it possible to provide a dynamic and personalized advertising experience tailored to the user.
[0333] "Video and audio information acquired from the user" refers to visual and auditory data collected through the user's video camera and microphone, which forms the basis for understanding the user's emotional state.
[0334] "Means for estimating emotional states" refers to algorithms or technologies that analyze collected video and audio information to identify the emotions expressed by the user.
[0335] "Means for automatically adjusting the actions and voice quality of a virtual character" refers to technology that changes the facial expressions and voice quality of a virtual character in real time according to the user's estimated emotional state.
[0336] "A means of generating a user emotion model and reflecting it in the representation of a virtual character" refers to a technology that creates a database of user emotion tendencies and then uses that information to create a model that reflects the actions and speech of a virtual character.
[0337] A "personalized advertising experience" is a method of providing users with an individually optimized experience by presenting advertisements whose content dynamically changes according to each user's preferences and emotions.
[0338] This invention is a system that analyzes a user's emotional state and provides a personalized advertising experience through a virtual character based on that analysis. This system processes the user's video and audio information in real time to detect emotions and reflect them in the virtual character's actions and voice quality.
[0339] First, the server acquires video and audio information from the user's device. For video analysis, visual data processing software such as "OpenCV" is used for face recognition and facial expression analysis. For audio analysis, audio analysis tools such as "Praat" are used to analyze the tone and tempo of the sound. From this information, a machine learning algorithm (e.g., TensorFlow) is applied to estimate the user's emotional state.
[0340] Next, based on the estimated emotions, the server uses an emotion engine to adjust the base model of the virtual character. This allows the virtual character's facial expressions and voice quality to change in real time according to the user's emotions. As a result, the user can feel as if they are interacting with a real person.
[0341] Users can customize the appearance and personality of their virtual characters through the provided interface. This allows for the creation of virtual characters with characteristics tailored to the user's needs and the company's advertising goals.
[0342] Ultimately, the refined virtual persona is distributed via the device to various media outlets such as television, internet advertising, and social media. This makes it possible to provide a customized advertising experience for each user, leading to the expectation of higher advertising effectiveness.
[0343] For example, if a user smiles, the server quickly picks up on that emotion and adjusts the virtual character to give a positive response such as, "That's good news!" This improves the quality of the ad interaction.
[0344] Examples of prompts include, "How should the virtual character's expression change when the user's facial expression changes?" and "How should the virtual character speak to a user who is feeling anxious?" By using these prompts, the AI model can learn to achieve more natural dialogue.
[0345] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0346] Step 1:
[0347] The server retrieves video and audio data from the user.
[0348] The input consists of video and audio information acquired in real time from the user's camera and microphone. The server receives this data as a stream and prepares it for analysis. Specifically, the data is sent to the server via an API.
[0349] Step 2:
[0350] The server analyzes video and audio data to estimate the user's emotional state.
[0351] The input consists of video and audio data acquired in Step 1. The server uses machine learning algorithms (e.g., TensorFlow) to perform analysis, using "OpenCV" for facial expression analysis and "Praat" for audio analysis. The output is an emotional state profile that represents the user's emotions. Specifically, the server extracts facial features and applies algorithms to measure the tone and tempo of the audio.
[0352] Step 3:
[0353] The server handles the adjustments of the virtual characters.
[0354] As input, the system receives an emotional state profile from Step 2. The server uses a generative AI model and leverages an emotion engine to refine the base model of the virtual character. The output is the refined virtual character data, which includes changes in facial expressions and voice. In terms of specific actions, the server modifies the virtual character's movements and voice quality in real time according to the emotional profile.
[0355] Step 4:
[0356] Users customize their virtual characters.
[0357] The input consists of customization information provided by the user through the operation screen. Users can set the appearance, personality, voice quality, etc., of a virtual character, specifying characteristics that match the company's advertising goals. The output is the user-customized settings information of the virtual character. Specifically, changes specified by the user are reflected on the server through interaction via the GUI.
[0358] Step 5:
[0359] The device distributes virtual characters to media platforms.
[0360] The input is data for a virtual person, which is adjusted in step 3 and customized in step 4. The device delivers this data to users through television, social media, online advertising, etc. The output is the user's visual and auditory experience. Specifically, the device is integrated with existing streaming services and advertising delivery systems to deliver the virtual person in real time.
[0361] (Application Example 2)
[0362] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0363] In today's world, simply presenting products is often insufficient for corporate advertising to effectively appeal to consumers. In particular, recognizing the emotional state of individual consumers and presenting the most appropriate advertisements accordingly is an effective approach. However, conventional systems have been inadequate in real-time emotion recognition and the dynamic adjustment of advertisements based on that recognition. Therefore, there is a need to build systems that provide a more personalized advertising experience based on consumer emotions.
[0364] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0365] In this invention, the server includes means for collecting data to recognize a person's emotional state in real time and dynamically adjust the advertising content according to that emotional state; means for analyzing the collected data to create a basic model of the virtual person's actions and voice; and means for adjusting the virtual person's actions and voice based on the emotion analysis data. This makes it possible to dynamically adjust the advertising content based on the emotions of each individual consumer and provide effective advertising in real time.
[0366] "A person's emotional state" refers to an individual's inner feelings and mood, as perceived from their facial expressions, voice, and other cues.
[0367] "Real-time recognition" refers to a process where information is processed as soon as it is generated, and results are obtained immediately.
[0368] "Dynamically adjusting ad content" means instantly changing the information presented as an ad according to the user's situation and emotions.
[0369] "Means of data collection" refers to any technical device or method for acquiring audio, image, or other sensory data.
[0370] A "foundation model" refers to a basic system or framework built to provide a foundation for a specific application.
[0371] "Emotional analysis data" refers to the results of data analysis used to identify a user's emotional state.
[0372] "Means for adjusting the behavior and voice of a virtual character" refers to a mechanism for changing the movements and speech patterns of a generated virtual character.
[0373] "Feedback data" refers to data collected from users' reactions and opinions to the information provided by the system.
[0374] A "media platform" is an online or offline technological infrastructure for distributing information and content.
[0375] A server plays a central role in implementing this invention. The server first collects the user's facial expressions and voice data through smart glasses or other devices. This data is analyzed in real time using emotion recognition software such as Google Cloud Vision or Amazon Rekognition to identify the user's emotional state. The analysis results are instantly reflected in the base model of the virtual person associated with each user.
[0376] On the server, the virtual character's actions and voice are adjusted based on analyzed emotion data. Using a generative AI model, the virtual character's behavior is designed to dynamically adjust the ad content according to the user's emotions. This virtual character is projected in real time onto smart glasses displays and other device screens, displaying ads in an interactive manner.
[0377] Furthermore, users can customize their virtual avatars through the interface, setting their appearance, voice, and behavior according to the company's marketing strategy. Feedback data is also collected and stored on the server for improvement.
[0378] For example, if a user wearing smart glasses is detected as being in a relaxed state, advertisements for calming relaxation products will be displayed with a bright and gentle voice. In this way, personalized advertising tailored to each user's emotional state is possible.
[0379] An example of a prompt for a generative AI model is, "Please think of the best advertisement to display when the user is relaxed."
[0380] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0381] Step 1:
[0382] The server collects video and audio data in real time from the user's smart glasses via the camera and microphone. This input data includes the user's facial expressions, voice tone, and tempo. This data forms the basis for analyzing the user's emotions.
[0383] Step 2:
[0384] The server inputs the collected video and audio data into emotion recognition software such as Google Cloud Vision or Amazon Rekognition to analyze the user's emotional state. This process uses machine learning algorithms to output emotion labels such as "happiness" and "stress" from the input data. This is then incorporated into the base model and used to adjust the virtual character.
[0385] Step 3:
[0386] The server uses a generative AI model to adjust the behavior and voice of the virtual character based on the analyzed emotional state. For example, if the user is relaxed, it generates a virtual character with a cheerful and calm voice and facial expression. This process also influences the selection of advertising content, leading to the output of optimal advertisements.
[0387] Step 4:
[0388] The device displays a customized virtual person and advertisement content on the smart glasses' display. In this step, the generated virtual person interacts with the user, creating a personalized advertising experience. Based on the output information from the server, the smart glasses play high-resolution video and audio.
[0389] Step 5:
[0390] Through the displayed advertising interface, users input settings for a virtual character and provide feedback on the ad content on their device. This feedback data is sent to the server and used to improve personalization in the future, contributing to overall system performance improvements.
[0391] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0392] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0393] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0394] [Third Embodiment]
[0395] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0396] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0397] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0398] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0399] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0400] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0401] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0402] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0403] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0404] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0405] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0406] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0407] This invention provides a system for using virtual characters in a company's advertising activities and a method for supporting advertising activities based on this system. The system aims to efficiently perform advanced data analysis, generate behavioral and voice models, customize them, and deploy them in advertising activities.
[0408] First, the server collects a wide range of data necessary for generating a virtual character. This includes image data, audio data, and cultural and market background data. This data forms the basis for determining the behavior and characteristics of the virtual character.
[0409] Next, the server analyzes the collected data to create a basic model of the virtual character's movements and voice. This allows for a highly accurate reproduction of a digital character with realistic human facial expressions, movements, and vocal qualities.
[0410] Subsequently, companies customize the virtual persona to create one that suits their brand image. Through the interface, users can adjust the virtual persona's appearance, voice tone, and even its behavior in specific scenarios. This customization process allows companies to easily design a unique virtual persona that perfectly fits their market strategy.
[0411] Once the virtual character creation is complete, the device is responsible for distribution to various media platforms. The generated digital data of the virtual character can be widely applied to television commercials, online videos, social media content, and more. This allows companies to efficiently conduct global advertising campaigns.
[0412] Furthermore, after running an advertising campaign, users can send feedback data to the server. This data includes viewership ratings, engagement, and social media reactions. Based on this feedback, the server can continuously improve the performance of the virtual character and optimize it in response to new market trends.
[0413] One example is when an international beverage company launches a new product and uses a virtual character that matches its youthful image. This virtual character can deliver speeches tailored to the local language and culture, effectively conveying a consistent brand message across various countries. This allows the company to reduce advertising campaign costs while conducting effective promotions.
[0414] The following describes the processing flow.
[0415] Step 1:
[0416] The server collects the data necessary for generating virtual characters from the internet and internal databases. Specifically, it gathers image data, audio clips, text information, and data on cultural background to build the foundation for the reliability and expressiveness of the virtual characters.
[0417] Step 2:
[0418] The server analyzes the collected data and creates a basic model of the virtual character's movements and voice. Machine learning algorithms are used to extract patterns from the collected images and audio, and based on these, a model is designed to reproduce natural movements and speech patterns.
[0419] Step 3:
[0420] Users customize virtual characters through the provided interface. They can select appearances, personalities, and voices that align with the company's brand image, and set the message the virtual character should convey.
[0421] Step 4:
[0422] The server generates the final virtual character based on the customization information specified by the user. The generation process utilizes advanced 3D rendering technology to create the customized model in real time and save it as digital data.
[0423] Step 5:
[0424] The device distributes the generated virtual character to various media platforms such as television, the internet, and social media. Because the distribution is done in a format optimized for each media platform, high-quality visual presentation is possible.
[0425] Step 6:
[0426] Users send feedback data to the server to check the effectiveness of their advertising campaigns. This feedback data includes viewership ratings, viewer engagement, and comment content, which is used to measure the campaign's effectiveness.
[0427] Step 7:
[0428] The server analyzes the received feedback data and makes adjustments to improve the performance of the virtual character. The results of this analysis are reflected in the creation of future virtual characters and advertising campaigns, forming the basis for achieving greater effectiveness.
[0429] (Example 1)
[0430] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0431] In modern advertising, there is a growing need for virtual characters that can flexibly adapt to target markets and cultures. However, traditional systems have struggled to efficiently and effectively generate, customize, distribute, and improve these virtual characters. Therefore, there is a need for methods that maximize advertising effectiveness while keeping costs down.
[0432] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0433] In this invention, the server includes means for collecting a wide range of information necessary for generating a virtual character, means for analyzing the collected information and creating the basic structure of the virtual character's movements and voice, and means for adjusting the virtual character to match the company's brand image. This makes it possible to efficiently generate a virtual character optimized for the target market and utilize it in various media.
[0434] A "virtual character" is a computer-generated character created using digital technology that can behave like a real person.
[0435] "Information" refers to a dataset containing image data, audio data, and cultural and market background data necessary for generating a virtual character.
[0436] "Device" refers to a combination of hardware and software used to collect, analyze, and process information according to a specific purpose.
[0437] "Analysis" is a data processing process performed to derive the basic structure of a virtual character's movements and voice from the collected data.
[0438] "Basic structure" refers to the fundamental frameworks and algorithms necessary to represent the actions and voices of virtual characters.
[0439] "Brand image" refers to the characteristic of a company that uses fictional characters to express its philosophy and values and appeal to consumers.
[0440] "Adjustment" is the process of modifying a virtual character's appearance, voice, and behavior to meet specific requirements.
[0441] This invention is a system aimed at generating and applying virtual characters in corporate advertising activities. The system aims to maximize advertising effectiveness by analyzing various data, creating customizable virtual characters, and distributing them across multiple media platforms.
[0442] The server collects a wide range of information necessary for generating virtual characters. This process involves using web scraping tools and database management systems to gather images, audio, and data related to cultural background. Specific software examples include tools like Beautiful Soup and Scrapy, which are used for data collection.
[0443] Next, the server analyzes the acquired information and uses machine learning algorithms to create the basic structure of the virtual character's movements and voice. Here, machine learning libraries such as TensorFlow and PyTorch are used to train the model. This generates a character that behaves like a real human.
[0444] Based on the generated foundational structure, users can customize the virtual character to match the company's brand image through a dedicated interface. The UI is designed to be user-friendly, allowing users to adjust appearance, voice tone, and behavior in specific scenarios using sliders and menus.
[0445] Subsequently, the device processes the customized virtual character's digital data and distributes it in the appropriate format to a wide variety of media, including television commercials, online video platforms, and social media. This process includes converting the data using video encoding tools such as FFmpeg and distributing it to media platforms using APIs.
[0446] Finally, users collect feedback on advertising activities, send it to the server, and use it to improve the performance of the virtual character. Data analysis tools such as Google Analytics are used to analyze the feedback data, which allows the virtual character to continuously improve in response to market trends.
[0447] A concrete example would be a company using this system to generate a virtual character that matches the image of a new product when promoting it, and then launching an advertising campaign in the international market. In this case, an example of a prompt might be, "Please generate a youthful and approachable character to support the new product launch campaign."
[0448] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0449] Step 1:
[0450] The server collects a wide range of information necessary for generating virtual characters. As input, it obtains relevant image data, audio data, and cultural / market background data from the internet and internal databases. This data is collected using web scraping tools (e.g., Beautiful Soup, Scrapy). As output, it provides the raw dataset necessary for creating the base model of the virtual character.
[0451] Step 2:
[0452] The server analyzes the collected data and generates the basic structure of the virtual character's movements and voice. The input is the dataset obtained in Step 1. For analysis, machine learning libraries such as TensorFlow and PyTorch are used to train image recognition models and speech synthesis models. As output, basic structure data of a virtual character with realistic movements and voice is generated.
[0453] Step 3:
[0454] The user customizes the virtual character to match the company's brand image using a dedicated interface. The input is the basic structure data generated in step 2. The user manipulates sliders and menus to adjust the virtual character's appearance, voice tone, and specific behaviors. The output is the configuration data of the customized virtual character.
[0455] Step 4:
[0456] The terminal converts the customized virtual person into a digital format suitable for various media formats and distributes it. The input is the configuration data created in step 3. Using a video encoding tool such as FFmpeg, the digital data of the virtual person is converted into a format suitable for television commercials, online videos, and social media content. The output is content in a format suitable for each media platform.
[0457] Step 5:
[0458] Users collect feedback data on their advertising activities and send it to the server. The input is the results of advertising activities, such as viewership and engagement. Data analysis tools such as Google Analytics are used for analysis, and performance evaluations and improvement suggestions for the virtual character are generated. The output is feedback data provided for future updates and improvements to the virtual character.
[0459] (Application Example 1)
[0460] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0461] In modern advertising, there is a growing demand for more personalized messaging and promotions tailored to users' interests and preferences. Traditional methods often result in static, generalized advertisements, making it difficult to efficiently deliver compelling messages to individual users. Furthermore, the inability to quickly utilize feedback from advertising campaigns leads to limited advertising effectiveness.
[0462] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0463] In this invention, the server includes means for collecting information to generate a virtual character to be used in a medium; means for analyzing the collected information to create a basic model of the virtual character's movements and voice; means for adjusting the virtual character based on the requirements of the medium; means for generating the final virtual character based on the adjusted settings; means for distributing the generated virtual character to multiple information distribution platforms; means for improving the performance of the virtual character based on feedback data; and means for customizing the virtual character based on user interests and effectively conveying messages through videos and interactive simulations. This makes it possible to provide advertisements optimized for each user and maximize the effectiveness of advertising activities.
[0464] A "medium" is a method or technical means for transmitting a message over a wide area.
[0465] A "virtual character" is a character created using digital technology that imitates human characteristics and behaviors.
[0466] "Information" refers to the data necessary for creating and customizing virtual characters, and is handled at each stage of collection, analysis, and application.
[0467] A "server" is a computer system that stores, analyzes, and processes data, and supports the process of creating virtual characters.
[0468] A "distribution platform" is an infrastructure for distributing content and providing information, and it plays a role in conveying messages from virtual characters to users.
[0469] "Feedback data" refers to evaluation information regarding the performance of advertising campaigns and virtual characters, and serves as a basis for improvement and adjustments.
[0470] "Customization" is the act of adjusting the appearance and behavior of a virtual character to suit specific requirements and preferences, and is an important process for meeting user expectations.
[0471] "Simulation" is a technology that recreates a virtual situation on a computer and predicts specific results or effects, providing an interactive experience for the user.
[0472] A "message" is the information and brand image delivered to users through advertising activities, and is the central content of the communication.
[0473] The system for implementing this invention streamlines the process of generating virtual characters and utilizing them in advertising activities. The server first collects information necessary to generate virtual characters to be used in the media, such as image data, audio data, and market background information. Based on this collected information, it builds a basic model to mimic the actions and voice of the virtual character. In this process, it uses TensorFlow, a machine learning platform, to perform data processing and calculations to reproduce the natural behavior and voice characteristics of the character.
[0474] The device provides the ability to customize virtual characters through user interaction. Users can intuitively adjust the appearance and voice tone of the virtual character through the interface. Using Unity, it is possible to render the virtual character's movements in real time and provide visual feedback. Specifically, by using the app and adjusting the virtual character to suit a particular scenario, a character optimized for on-the-spot communication can be generated.
[0475] The digital data of virtual characters generated through the distribution platform is widely deployed in online advertising, social media content, and other areas. This process makes it possible to efficiently implement global advertising campaigns.
[0476] Furthermore, users send feedback to the server regarding the performance of their advertising campaigns. This feedback includes viewership ratings, engagement, and social media reactions, and the analysis results are used to improve the performance of virtual characters and optimize them in response to market trends.
[0477] As a concrete example, a prompt for generating a virtual character who introduces casual fashion for women in their 20s could be: "Please generate a virtual character who introduces popular casual fashion items for women in their 20s. The character should have a bright expression, greet in a friendly tone, and clearly explain the features of the products."
[0478] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0479] Step 1:
[0480] The server collects the information necessary to generate virtual characters used in the media. This includes image data, audio data, and cultural and market background information from multiple data sources. The input is this diverse set of data, and the output is a basic dataset for virtual character generation. Specifically, it automates data acquisition through web scraping and APIs.
[0481] Step 2:
[0482] The server analyzes the collected information and creates a basic model of the virtual character's movements and speech. It processes the data using TensorFlow and builds a machine learning model. The input is the basic dataset, and the output is the virtual character's movements and speech model. Specifically, it performs data cleaning, feature selection, and model training.
[0483] Step 3:
[0484] Users customize their virtual character through the device's interface. This includes adjusting appearance, voice tone, and behavior. Input is configuration information from the base model and user interface, and output is the customized virtual character model. Specifically, intuitive configuration changes are possible via the UI.
[0485] Step 4:
[0486] The device uses Unity to render virtual characters in real time and provide visual feedback. The input is a customized virtual character model, and the output is an interactive character displayed to the user. Specifically, it handles the rendering of 3D models and user interaction.
[0487] Step 5:
[0488] The server sends the generated virtual character's digital data to the distribution platform for distribution as online advertisements and social media content. The input is the completed virtual character model data, and the output is its display on the distribution platform. Specifically, it configures content distribution settings through the APIs of each platform.
[0489] Step 6:
[0490] Users send feedback to the server regarding the performance of their advertising campaigns. This feedback includes data such as viewership and engagement; the input is the campaign results, and the output is an analytical report for improvement. Specifically, data analysis software is used to evaluate the results and propose improvements.
[0491] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0492] This invention relates to a virtual character generation system that incorporates an emotion engine, with the aim of achieving more personalized effects in corporate advertising activities. This system has the ability to recognize the user's emotions in real time and automatically adjust the behavior and voice expression of the virtual character accordingly.
[0493] First, as an embodiment, the server collects and analyzes video and audio data for emotion recognition. Specifically, it measures the user's facial expressions, voice tone, and tempo, and analyzes this data using a machine learning algorithm to identify the user's emotional state. The collected data is used to build an emotion model, which is then reflected in the behavior of the virtual character.
[0494] Next, the server uses an emotion engine to make adjustments to the virtual character's base model that correspond to the user's emotions. These adjustments allow the virtual character to display changes in facial expressions and voice quality in accordance with the user's emotions, enabling more realistic communication. For example, when the user expresses happiness, the virtual character is adjusted to speak with a brighter expression and a more cheerful tone.
[0495] Furthermore, users can further customize the virtual character to match the characteristics of the company's brand. They can use the interface to set the appearance and personality of the virtual character, providing a personalized user experience tailored to the company's marketing objectives.
[0496] The generated virtual characters are distributed to various media platforms via the device. When distributed through television, online advertising, social media, etc., the aforementioned emotional response is utilized in interactions with actual users. This makes it possible to further enhance the effectiveness of advertising.
[0497] One possible use case is an international cosmetics company using this system to promote a new product. In this scenario, when a target user interacts through the camera, their emotions, such as interest or surprise, are identified, and a virtual character responds by providing product information and messages accordingly. This provides a more personalized advertising experience, increasing the user's interest in the product and their willingness to purchase it.
[0498] The following describes the processing flow.
[0499] Step 1:
[0500] The server collects the data necessary to recognize the user's emotions. Specifically, it acquires the user's facial expression and voice data through cameras and microphones, and inputs this data into the emotion analysis system.
[0501] Step 2:
[0502] The server uses machine learning algorithms to analyze the user's emotions based on acquired facial and voice data. Here, emotions are identified using elements such as smiles, the degree of frowning, and the tone and rhythm of the voice, and the emotional state is recognized in real time.
[0503] Step 3:
[0504] The server dynamically adjusts the virtual character's behavior and voice expression based on the results of emotion analysis. For example, if the user indicates feelings of joy, the server instructs the virtual character to respond with a cheerful expression and a friendly tone.
[0505] Step 4:
[0506] Users customize their virtual avatars through the interface. They can set the appearance and style of their virtual avatars to match specific advertising campaigns, providing an experience tailored to their individual preferences.
[0507] Step 5:
[0508] The server generates a final model of the virtual character based on customization and emotion recognition, and stores it as digital data suitable for use. This enables consistent representation of the virtual character across various media platforms.
[0509] Step 6:
[0510] The device distributes the generated virtual character to multiple media platforms, including television, online advertising, and social media. The virtual character is used as part of digital content, enabling effective communication that reflects the user's emotional responses.
[0511] Step 7:
[0512] Users provide feedback on the performance of advertising campaigns and user response data, and the server analyzes this data to further improve the performance of the virtual character and emotion engine. Through this process, the ability to more accurately target the audience in the next campaign is enhanced.
[0513] (Example 2)
[0514] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0515] Traditional advertising methods have struggled to deliver personalized ads that accurately capture the emotions of target users, resulting in limited advertising effectiveness. In particular, providing a consistent customer experience across diverse media platforms is difficult, and the inability to utilize real-time feedback that reflects user experience is a significant problem.
[0516] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0517] In this invention, the server includes means for analyzing video and audio information acquired from the user to estimate the user's emotional state, means for automatically adjusting the actions and voice quality of the virtual character based on the estimated emotional state, and means for generating a user's emotional model and reflecting it in the virtual character's expression. This makes it possible to provide a dynamic and personalized advertising experience tailored to the user.
[0518] "Video and audio information acquired from the user" refers to visual and auditory data collected through the user's video camera and microphone, which forms the basis for understanding the user's emotional state.
[0519] "Means for estimating emotional states" refers to algorithms or technologies that analyze collected video and audio information to identify the emotions expressed by the user.
[0520] "Means for automatically adjusting the actions and voice quality of a virtual character" refers to technology that changes the facial expressions and voice quality of a virtual character in real time according to the user's estimated emotional state.
[0521] "A means of generating a user emotion model and reflecting it in the representation of a virtual character" refers to a technology that creates a database of user emotion tendencies and then uses that information to create a model that reflects the actions and speech of a virtual character.
[0522] A "personalized advertising experience" is a method of providing users with an individually optimized experience by presenting advertisements whose content dynamically changes according to each user's preferences and emotions.
[0523] This invention is a system that analyzes a user's emotional state and provides a personalized advertising experience through a virtual character based on that analysis. This system processes the user's video and audio information in real time to detect emotions and reflect them in the virtual character's actions and voice quality.
[0524] First, the server acquires video and audio information from the user's device. For video analysis, visual data processing software such as "OpenCV" is used for face recognition and facial expression analysis. For audio analysis, audio analysis tools such as "Praat" are used to analyze the tone and tempo of the sound. From this information, a machine learning algorithm (e.g., TensorFlow) is applied to estimate the user's emotional state.
[0525] Next, based on the estimated emotions, the server uses an emotion engine to adjust the base model of the virtual character. This allows the virtual character's facial expressions and voice quality to change in real time according to the user's emotions. As a result, the user can feel as if they are interacting with a real person.
[0526] Users can customize the appearance and personality of their virtual characters through the provided interface. This allows for the creation of virtual characters with characteristics tailored to the user's needs and the company's advertising goals.
[0527] Ultimately, the refined virtual persona is distributed via the device to various media outlets such as television, internet advertising, and social media. This makes it possible to provide a customized advertising experience for each user, leading to the expectation of higher advertising effectiveness.
[0528] For example, if a user smiles, the server quickly picks up on that emotion and adjusts the virtual character to give a positive response such as, "That's good news!" This improves the quality of the ad interaction.
[0529] Examples of prompts include, "How should the virtual character's expression change when the user's facial expression changes?" and "How should the virtual character speak to a user who is feeling anxious?" By using these prompts, the AI model can learn to achieve more natural dialogue.
[0530] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0531] Step 1:
[0532] The server retrieves video and audio data from the user.
[0533] The input consists of video and audio information acquired in real time from the user's camera and microphone. The server receives this data as a stream and prepares it for analysis. Specifically, the data is sent to the server via an API.
[0534] Step 2:
[0535] The server analyzes video and audio data to estimate the user's emotional state.
[0536] The input consists of video and audio data acquired in Step 1. The server uses machine learning algorithms (e.g., TensorFlow) to perform analysis, using "OpenCV" for facial expression analysis and "Praat" for audio analysis. The output is an emotional state profile that represents the user's emotions. Specifically, the server extracts facial features and applies algorithms to measure the tone and tempo of the audio.
[0537] Step 3:
[0538] The server handles the adjustments of the virtual characters.
[0539] As input, the system receives an emotional state profile from Step 2. The server uses a generative AI model and leverages an emotion engine to refine the base model of the virtual character. The output is the refined virtual character data, which includes changes in facial expressions and voice. In terms of specific actions, the server modifies the virtual character's movements and voice quality in real time according to the emotional profile.
[0540] Step 4:
[0541] Users customize their virtual characters.
[0542] The input consists of customization information provided by the user through the operation screen. Users can set the appearance, personality, voice quality, etc., of a virtual character, specifying characteristics that match the company's advertising goals. The output is the user-customized settings information of the virtual character. Specifically, changes specified by the user are reflected on the server through interaction via the GUI.
[0543] Step 5:
[0544] The device distributes virtual characters to media platforms.
[0545] The input is data for a virtual person, which is adjusted in step 3 and customized in step 4. The device delivers this data to users through television, social media, online advertising, etc. The output is the user's visual and auditory experience. Specifically, the device is integrated with existing streaming services and advertising delivery systems to deliver the virtual person in real time.
[0546] (Application Example 2)
[0547] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0548] In today's world, simply presenting products is often insufficient for corporate advertising to effectively appeal to consumers. In particular, recognizing the emotional state of individual consumers and presenting the most appropriate advertisements accordingly is an effective approach. However, conventional systems have been inadequate in real-time emotion recognition and the dynamic adjustment of advertisements based on that recognition. Therefore, there is a need to build systems that provide a more personalized advertising experience based on consumer emotions.
[0549] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0550] In this invention, the server includes means for collecting data to recognize a person's emotional state in real time and dynamically adjust the advertising content according to that emotional state; means for analyzing the collected data to create a basic model of the virtual person's actions and voice; and means for adjusting the virtual person's actions and voice based on the emotion analysis data. This makes it possible to dynamically adjust the advertising content based on the emotions of each individual consumer and provide effective advertising in real time.
[0551] "A person's emotional state" refers to an individual's inner feelings and mood, as perceived from their facial expressions, voice, and other cues.
[0552] "Real-time recognition" refers to a process where information is processed as soon as it is generated, and results are obtained immediately.
[0553] "Dynamically adjusting ad content" means instantly changing the information presented as an ad according to the user's situation and emotions.
[0554] "Means of data collection" refers to any technical device or method for acquiring audio, image, or other sensory data.
[0555] A "foundation model" refers to a basic system or framework built to provide a foundation for a specific application.
[0556] "Emotional analysis data" refers to the results of data analysis used to identify a user's emotional state.
[0557] "Means for adjusting the behavior and voice of a virtual character" refers to a mechanism for changing the movements and speech patterns of a generated virtual character.
[0558] "Feedback data" refers to data collected from users' reactions and opinions to the information provided by the system.
[0559] A "media platform" is an online or offline technological infrastructure for distributing information and content.
[0560] A server plays a central role in implementing this invention. The server first collects the user's facial expressions and voice data through smart glasses or other devices. This data is analyzed in real time using emotion recognition software such as Google Cloud Vision or Amazon Rekognition to identify the user's emotional state. The analysis results are instantly reflected in the base model of the virtual person associated with each user.
[0561] On the server, the virtual character's actions and voice are adjusted based on analyzed emotion data. Using a generative AI model, the virtual character's behavior is designed to dynamically adjust the ad content according to the user's emotions. This virtual character is projected in real time onto smart glasses displays and other device screens, displaying ads in an interactive manner.
[0562] Furthermore, users can customize their virtual avatars through the interface, setting their appearance, voice, and behavior according to the company's marketing strategy. Feedback data is also collected and stored on the server for improvement.
[0563] For example, if a user wearing smart glasses is detected as being in a relaxed state, advertisements for calming relaxation products will be displayed with a bright and gentle voice. In this way, personalized advertising tailored to each user's emotional state is possible.
[0564] An example of a prompt for a generative AI model is, "Please think of the best advertisement to display when the user is relaxed."
[0565] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0566] Step 1:
[0567] The server collects video and audio data in real time from the user's smart glasses via the camera and microphone. This input data includes the user's facial expressions, voice tone, and tempo. This data forms the basis for analyzing the user's emotions.
[0568] Step 2:
[0569] The server inputs the collected video and audio data into emotion recognition software such as Google Cloud Vision or Amazon Rekognition to analyze the user's emotional state. This process uses machine learning algorithms to output emotion labels such as "happiness" and "stress" from the input data. This is then incorporated into the base model and used to adjust the virtual character.
[0570] Step 3:
[0571] The server uses a generative AI model to adjust the behavior and voice of the virtual character based on the analyzed emotional state. For example, if the user is relaxed, it generates a virtual character with a cheerful and calm voice and facial expression. This process also influences the selection of advertising content, leading to the output of optimal advertisements.
[0572] Step 4:
[0573] The device displays a customized virtual person and advertisement content on the smart glasses' display. In this step, the generated virtual person interacts with the user, creating a personalized advertising experience. Based on the output information from the server, the smart glasses play high-resolution video and audio.
[0574] Step 5:
[0575] Through the displayed advertising interface, users input settings for a virtual character and provide feedback on the ad content on their device. This feedback data is sent to the server and used to improve personalization in the future, contributing to overall system performance improvements.
[0576] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0577] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0578] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0579] [Fourth Embodiment]
[0580] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0581] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0582] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0583] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0584] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0585] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0586] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0587] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0588] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0589] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0590] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0591] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0592] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0593] This invention provides a system for using virtual characters in a company's advertising activities and a method for supporting advertising activities based on this system. The system aims to efficiently perform advanced data analysis, generate behavioral and voice models, customize them, and deploy them in advertising activities.
[0594] First, the server collects a wide range of data necessary for generating a virtual character. This includes image data, audio data, and cultural and market background data. This data forms the basis for determining the behavior and characteristics of the virtual character.
[0595] Next, the server analyzes the collected data to create a basic model of the virtual character's movements and voice. This allows for a highly accurate reproduction of a digital character with realistic human facial expressions, movements, and vocal qualities.
[0596] Subsequently, companies customize the virtual persona to create one that suits their brand image. Through the interface, users can adjust the virtual persona's appearance, voice tone, and even its behavior in specific scenarios. This customization process allows companies to easily design a unique virtual persona that perfectly fits their market strategy.
[0597] Once the virtual character creation is complete, the device is responsible for distribution to various media platforms. The generated digital data of the virtual character can be widely applied to television commercials, online videos, social media content, and more. This allows companies to efficiently conduct global advertising campaigns.
[0598] Furthermore, after running an advertising campaign, users can send feedback data to the server. This data includes viewership ratings, engagement, and social media reactions. Based on this feedback, the server can continuously improve the performance of the virtual character and optimize it in response to new market trends.
[0599] One example is when an international beverage company launches a new product and uses a virtual character that matches its youthful image. This virtual character can deliver speeches tailored to the local language and culture, effectively conveying a consistent brand message across various countries. This allows the company to reduce advertising campaign costs while conducting effective promotions.
[0600] The following describes the processing flow.
[0601] Step 1:
[0602] The server collects the data necessary for generating virtual characters from the internet and internal databases. Specifically, it gathers image data, audio clips, text information, and data on cultural background to build the foundation for the reliability and expressiveness of the virtual characters.
[0603] Step 2:
[0604] The server analyzes the collected data and creates a basic model of the virtual character's movements and voice. Machine learning algorithms are used to extract patterns from the collected images and audio, and based on these, a model is designed to reproduce natural movements and speech patterns.
[0605] Step 3:
[0606] Users customize virtual characters through the provided interface. They can select appearances, personalities, and voices that align with the company's brand image, and set the message the virtual character should convey.
[0607] Step 4:
[0608] The server generates the final virtual character based on the customization information specified by the user. The generation process utilizes advanced 3D rendering technology to create the customized model in real time and save it as digital data.
[0609] Step 5:
[0610] The device distributes the generated virtual character to various media platforms such as television, the internet, and social media. Because the distribution is done in a format optimized for each media platform, high-quality visual presentation is possible.
[0611] Step 6:
[0612] Users send feedback data to the server to check the effectiveness of their advertising campaigns. This feedback data includes viewership ratings, viewer engagement, and comment content, which is used to measure the campaign's effectiveness.
[0613] Step 7:
[0614] The server analyzes the received feedback data and makes adjustments to improve the performance of the virtual character. The results of this analysis are reflected in the creation of future virtual characters and advertising campaigns, forming the basis for achieving greater effectiveness.
[0615] (Example 1)
[0616] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0617] In modern advertising, there is a growing need for virtual characters that can flexibly adapt to target markets and cultures. However, traditional systems have struggled to efficiently and effectively generate, customize, distribute, and improve these virtual characters. Therefore, there is a need for methods that maximize advertising effectiveness while keeping costs down.
[0618] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0619] In this invention, the server includes means for collecting a wide range of information necessary for generating a virtual character, means for analyzing the collected information and creating the basic structure of the virtual character's movements and voice, and means for adjusting the virtual character to match the company's brand image. This makes it possible to efficiently generate a virtual character optimized for the target market and utilize it in various media.
[0620] A "virtual character" is a computer-generated character created using digital technology that can behave like a real person.
[0621] "Information" refers to a dataset containing image data, audio data, and cultural and market background data necessary for generating a virtual character.
[0622] "Device" refers to a combination of hardware and software used to collect, analyze, and process information according to a specific purpose.
[0623] "Analysis" is a data processing process performed to derive the basic structure of a virtual character's movements and voice from the collected data.
[0624] "Basic structure" refers to the fundamental frameworks and algorithms necessary to represent the actions and voices of virtual characters.
[0625] "Brand image" refers to the characteristic of a company that uses fictional characters to express its philosophy and values and appeal to consumers.
[0626] "Adjustment" is the process of modifying a virtual character's appearance, voice, and behavior to meet specific requirements.
[0627] This invention is a system aimed at generating and applying virtual characters in corporate advertising activities. The system aims to maximize advertising effectiveness by analyzing various data, creating customizable virtual characters, and distributing them across multiple media platforms.
[0628] The server collects a wide range of information necessary for generating virtual characters. This process involves using web scraping tools and database management systems to gather images, audio, and data related to cultural background. Specific software examples include tools like Beautiful Soup and Scrapy, which are used for data collection.
[0629] Next, the server analyzes the acquired information and uses machine learning algorithms to create the basic structure of the virtual character's movements and voice. Here, machine learning libraries such as TensorFlow and PyTorch are used to train the model. This generates a character that behaves like a real human.
[0630] Based on the generated foundational structure, users can customize the virtual character to match the company's brand image through a dedicated interface. The UI is designed to be user-friendly, allowing users to adjust appearance, voice tone, and behavior in specific scenarios using sliders and menus.
[0631] Subsequently, the device processes the customized virtual character's digital data and distributes it in the appropriate format to a wide variety of media, including television commercials, online video platforms, and social media. This process includes converting the data using video encoding tools such as FFmpeg and distributing it to media platforms using APIs.
[0632] Finally, users collect feedback on advertising activities, send it to the server, and use it to improve the performance of the virtual character. Data analysis tools such as Google Analytics are used to analyze the feedback data, which allows the virtual character to continuously improve in response to market trends.
[0633] A concrete example would be a company using this system to generate a virtual character that matches the image of a new product when promoting it, and then launching an advertising campaign in the international market. In this case, an example of a prompt might be, "Please generate a youthful and approachable character to support the new product launch campaign."
[0634] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0635] Step 1:
[0636] The server collects a wide range of information necessary for generating virtual characters. As input, it obtains relevant image data, audio data, and cultural / market background data from the internet and internal databases. This data is collected using web scraping tools (e.g., Beautiful Soup, Scrapy). As output, it provides the raw dataset necessary for creating the base model of the virtual character.
[0637] Step 2:
[0638] The server analyzes the collected data and generates the basic structure of the virtual character's movements and voice. The input is the dataset obtained in Step 1. For analysis, machine learning libraries such as TensorFlow and PyTorch are used to train image recognition models and speech synthesis models. As output, basic structure data of a virtual character with realistic movements and voice is generated.
[0639] Step 3:
[0640] The user customizes the virtual character to match the company's brand image using a dedicated interface. The input is the basic structure data generated in step 2. The user manipulates sliders and menus to adjust the virtual character's appearance, voice tone, and specific behaviors. The output is the configuration data of the customized virtual character.
[0641] Step 4:
[0642] The terminal converts the customized virtual person into a digital format suitable for various media formats and distributes it. The input is the configuration data created in step 3. Using a video encoding tool such as FFmpeg, the digital data of the virtual person is converted into a format suitable for television commercials, online videos, and social media content. The output is content in a format suitable for each media platform.
[0643] Step 5:
[0644] Users collect feedback data on their advertising activities and send it to the server. The input is the results of advertising activities, such as viewership and engagement. Data analysis tools such as Google Analytics are used for analysis, and performance evaluations and improvement suggestions for the virtual character are generated. The output is feedback data provided for future updates and improvements to the virtual character.
[0645] (Application Example 1)
[0646] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0647] In modern advertising, there is a growing demand for more personalized messaging and promotions tailored to users' interests and preferences. Traditional methods often result in static, generalized advertisements, making it difficult to efficiently deliver compelling messages to individual users. Furthermore, the inability to quickly utilize feedback from advertising campaigns leads to limited advertising effectiveness.
[0648] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0649] In this invention, the server includes means for collecting information to generate a virtual character to be used in a medium; means for analyzing the collected information to create a basic model of the virtual character's movements and voice; means for adjusting the virtual character based on the requirements of the medium; means for generating the final virtual character based on the adjusted settings; means for distributing the generated virtual character to multiple information distribution platforms; means for improving the performance of the virtual character based on feedback data; and means for customizing the virtual character based on user interests and effectively conveying messages through videos and interactive simulations. This makes it possible to provide advertisements optimized for each user and maximize the effectiveness of advertising activities.
[0650] A "medium" is a method or technical means for transmitting a message over a wide area.
[0651] A "virtual character" is a character created using digital technology that imitates human characteristics and behaviors.
[0652] "Information" refers to the data necessary for creating and customizing virtual characters, and is handled at each stage of collection, analysis, and application.
[0653] A "server" is a computer system that stores, analyzes, and processes data, and supports the process of creating virtual characters.
[0654] A "distribution platform" is an infrastructure for distributing content and providing information, and it plays a role in conveying messages from virtual characters to users.
[0655] "Feedback data" refers to evaluation information regarding the performance of advertising campaigns and virtual characters, and serves as a basis for improvement and adjustments.
[0656] "Customization" is the act of adjusting the appearance and behavior of a virtual character to suit specific requirements and preferences, and is an important process for meeting user expectations.
[0657] "Simulation" is a technology that recreates a virtual situation on a computer and predicts specific results or effects, providing an interactive experience for the user.
[0658] A "message" is the information and brand image delivered to users through advertising activities, and is the central content of the communication.
[0659] The system for implementing this invention streamlines the process of generating virtual characters and utilizing them in advertising activities. The server first collects information necessary to generate virtual characters to be used in the media, such as image data, audio data, and market background information. Based on this collected information, it builds a basic model to mimic the actions and voice of the virtual character. In this process, it uses TensorFlow, a machine learning platform, to perform data processing and calculations to reproduce the natural behavior and voice characteristics of the character.
[0660] The device provides the ability to customize virtual characters through user interaction. Users can intuitively adjust the appearance and voice tone of the virtual character through the interface. Using Unity, it is possible to render the virtual character's movements in real time and provide visual feedback. Specifically, by using the app and adjusting the virtual character to suit a particular scenario, a character optimized for on-the-spot communication can be generated.
[0661] The digital data of virtual characters generated through the distribution platform is widely deployed in online advertising, social media content, and other areas. This process makes it possible to efficiently implement global advertising campaigns.
[0662] Furthermore, users send feedback to the server regarding the performance of their advertising campaigns. This feedback includes viewership ratings, engagement, and social media reactions, and the analysis results are used to improve the performance of virtual characters and optimize them in response to market trends.
[0663] As a concrete example, a prompt for generating a virtual character who introduces casual fashion for women in their 20s could be: "Please generate a virtual character who introduces popular casual fashion items for women in their 20s. The character should have a bright expression, greet in a friendly tone, and clearly explain the features of the products."
[0664] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0665] Step 1:
[0666] The server collects the information necessary to generate virtual characters used in the media. This includes image data, audio data, and cultural and market background information from multiple data sources. The input is this diverse set of data, and the output is a basic dataset for virtual character generation. Specifically, it automates data acquisition through web scraping and APIs.
[0667] Step 2:
[0668] The server analyzes the collected information and creates a basic model of the virtual character's movements and speech. It processes the data using TensorFlow and builds a machine learning model. The input is the basic dataset, and the output is the virtual character's movements and speech model. Specifically, it performs data cleaning, feature selection, and model training.
[0669] Step 3:
[0670] Users customize their virtual character through the device's interface. This includes adjusting appearance, voice tone, and behavior. Input is configuration information from the base model and user interface, and output is the customized virtual character model. Specifically, intuitive configuration changes are possible via the UI.
[0671] Step 4:
[0672] The device uses Unity to render virtual characters in real time and provide visual feedback. The input is a customized virtual character model, and the output is an interactive character displayed to the user. Specifically, it handles the rendering of 3D models and user interaction.
[0673] Step 5:
[0674] The server sends the generated virtual character's digital data to the distribution platform for distribution as online advertisements and social media content. The input is the completed virtual character model data, and the output is its display on the distribution platform. Specifically, it configures content distribution settings through the APIs of each platform.
[0675] Step 6:
[0676] Users send feedback to the server regarding the performance of their advertising campaigns. This feedback includes data such as viewership and engagement; the input is the campaign results, and the output is an analytical report for improvement. Specifically, data analysis software is used to evaluate the results and propose improvements.
[0677] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0678] This invention relates to a virtual character generation system that incorporates an emotion engine, with the aim of achieving more personalized effects in corporate advertising activities. This system has the ability to recognize the user's emotions in real time and automatically adjust the behavior and voice expression of the virtual character accordingly.
[0679] First, as an embodiment, the server collects and analyzes video and audio data for emotion recognition. Specifically, it measures the user's facial expressions, voice tone, and tempo, and analyzes this data using a machine learning algorithm to identify the user's emotional state. The collected data is used to build an emotion model, which is then reflected in the behavior of the virtual character.
[0680] Next, the server uses an emotion engine to make adjustments to the virtual character's base model that correspond to the user's emotions. These adjustments allow the virtual character to display changes in facial expressions and voice quality in accordance with the user's emotions, enabling more realistic communication. For example, when the user expresses happiness, the virtual character is adjusted to speak with a brighter expression and a more cheerful tone.
[0681] Furthermore, users can further customize the virtual character to match the characteristics of the company's brand. They can use the interface to set the appearance and personality of the virtual character, providing a personalized user experience tailored to the company's marketing objectives.
[0682] The generated virtual characters are distributed to various media platforms via the device. When distributed through television, online advertising, social media, etc., the aforementioned emotional response is utilized in interactions with actual users. This makes it possible to further enhance the effectiveness of advertising.
[0683] One possible use case is an international cosmetics company using this system to promote a new product. In this scenario, when a target user interacts through the camera, their emotions, such as interest or surprise, are identified, and a virtual character responds by providing product information and messages accordingly. This provides a more personalized advertising experience, increasing the user's interest in the product and their willingness to purchase it.
[0684] The following describes the processing flow.
[0685] Step 1:
[0686] The server collects the data necessary to recognize the user's emotions. Specifically, it acquires the user's facial expression and voice data through cameras and microphones, and inputs this data into the emotion analysis system.
[0687] Step 2:
[0688] The server uses machine learning algorithms to analyze the user's emotions based on acquired facial and voice data. Here, emotions are identified using elements such as smiles, the degree of frowning, and the tone and rhythm of the voice, and the emotional state is recognized in real time.
[0689] Step 3:
[0690] The server dynamically adjusts the virtual character's behavior and voice expression based on the results of emotion analysis. For example, if the user indicates feelings of joy, the server instructs the virtual character to respond with a cheerful expression and a friendly tone.
[0691] Step 4:
[0692] Users customize their virtual avatars through the interface. They can set the appearance and style of their virtual avatars to match specific advertising campaigns, providing an experience tailored to their individual preferences.
[0693] Step 5:
[0694] The server generates a final model of the virtual character based on customization and emotion recognition, and stores it as digital data suitable for use. This enables consistent representation of the virtual character across various media platforms.
[0695] Step 6:
[0696] The device distributes the generated virtual character to multiple media platforms, including television, online advertising, and social media. The virtual character is used as part of digital content, enabling effective communication that reflects the user's emotional responses.
[0697] Step 7:
[0698] Users provide feedback on the performance of advertising campaigns and user response data, and the server analyzes this data to further improve the performance of the virtual character and emotion engine. Through this process, the ability to more accurately target the audience in the next campaign is enhanced.
[0699] (Example 2)
[0700] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0701] Traditional advertising methods have struggled to deliver personalized ads that accurately capture the emotions of target users, resulting in limited advertising effectiveness. In particular, providing a consistent customer experience across diverse media platforms is difficult, and the inability to utilize real-time feedback that reflects user experience is a significant problem.
[0702] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0703] In this invention, the server includes means for analyzing video and audio information acquired from the user to estimate the user's emotional state, means for automatically adjusting the actions and voice quality of the virtual character based on the estimated emotional state, and means for generating a user's emotional model and reflecting it in the virtual character's expression. This makes it possible to provide a dynamic and personalized advertising experience tailored to the user.
[0704] "Video and audio information acquired from the user" refers to visual and auditory data collected through the user's video camera and microphone, which forms the basis for understanding the user's emotional state.
[0705] "Means for estimating emotional states" refers to algorithms or technologies that analyze collected video and audio information to identify the emotions expressed by the user.
[0706] "Means for automatically adjusting the actions and voice quality of a virtual character" refers to technology that changes the facial expressions and voice quality of a virtual character in real time according to the user's estimated emotional state.
[0707] "A means of generating a user emotion model and reflecting it in the representation of a virtual character" refers to a technology that creates a database of user emotion tendencies and then uses that information to create a model that reflects the actions and speech of a virtual character.
[0708] A "personalized advertising experience" is a method of providing users with an individually optimized experience by presenting advertisements whose content dynamically changes according to each user's preferences and emotions.
[0709] This invention is a system that analyzes a user's emotional state and provides a personalized advertising experience through a virtual character based on that analysis. This system processes the user's video and audio information in real time to detect emotions and reflect them in the virtual character's actions and voice quality.
[0710] First, the server acquires video and audio information from the user's device. For video analysis, visual data processing software such as "OpenCV" is used for face recognition and facial expression analysis. For audio analysis, audio analysis tools such as "Praat" are used to analyze the tone and tempo of the sound. From this information, a machine learning algorithm (e.g., TensorFlow) is applied to estimate the user's emotional state.
[0711] Next, based on the estimated emotions, the server uses an emotion engine to adjust the base model of the virtual character. This allows the virtual character's facial expressions and voice quality to change in real time according to the user's emotions. As a result, the user can feel as if they are interacting with a real person.
[0712] Users can customize the appearance and personality of their virtual characters through the provided interface. This allows for the creation of virtual characters with characteristics tailored to the user's needs and the company's advertising goals.
[0713] Ultimately, the refined virtual persona is distributed via the device to various media outlets such as television, internet advertising, and social media. This makes it possible to provide a customized advertising experience for each user, leading to the expectation of higher advertising effectiveness.
[0714] For example, if a user smiles, the server quickly picks up on that emotion and adjusts the virtual character to give a positive response such as, "That's good news!" This improves the quality of the ad interaction.
[0715] Examples of prompts include, "How should the virtual character's expression change when the user's facial expression changes?" and "How should the virtual character speak to a user who is feeling anxious?" By using these prompts, the AI model can learn to achieve more natural dialogue.
[0716] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0717] Step 1:
[0718] The server retrieves video and audio data from the user.
[0719] The input consists of video and audio information acquired in real time from the user's camera and microphone. The server receives this data as a stream and prepares it for analysis. Specifically, the data is sent to the server via an API.
[0720] Step 2:
[0721] The server analyzes video and audio data to estimate the user's emotional state.
[0722] The input consists of video and audio data acquired in Step 1. The server uses machine learning algorithms (e.g., TensorFlow) to perform analysis, using "OpenCV" for facial expression analysis and "Praat" for audio analysis. The output is an emotional state profile that represents the user's emotions. Specifically, the server extracts facial features and applies algorithms to measure the tone and tempo of the audio.
[0723] Step 3:
[0724] The server handles the adjustments of the virtual characters.
[0725] As input, the system receives an emotional state profile from Step 2. The server uses a generative AI model and leverages an emotion engine to refine the base model of the virtual character. The output is the refined virtual character data, which includes changes in facial expressions and voice. In terms of specific actions, the server modifies the virtual character's movements and voice quality in real time according to the emotional profile.
[0726] Step 4:
[0727] Users customize their virtual characters.
[0728] The input consists of customization information provided by the user through the operation screen. Users can set the appearance, personality, voice quality, etc., of a virtual character, specifying characteristics that match the company's advertising goals. The output is the user-customized settings information of the virtual character. Specifically, changes specified by the user are reflected on the server through interaction via the GUI.
[0729] Step 5:
[0730] The device distributes virtual characters to media platforms.
[0731] The input is data for a virtual person, which is adjusted in step 3 and customized in step 4. The device delivers this data to users through television, social media, online advertising, etc. The output is the user's visual and auditory experience. Specifically, the device is integrated with existing streaming services and advertising delivery systems to deliver the virtual person in real time.
[0732] (Application Example 2)
[0733] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0734] In today's world, simply presenting products is often insufficient for corporate advertising to effectively appeal to consumers. In particular, recognizing the emotional state of individual consumers and presenting the most appropriate advertisements accordingly is an effective approach. However, conventional systems have been inadequate in real-time emotion recognition and the dynamic adjustment of advertisements based on that recognition. Therefore, there is a need to build systems that provide a more personalized advertising experience based on consumer emotions.
[0735] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0736] In this invention, the server includes means for collecting data to recognize a person's emotional state in real time and dynamically adjust the advertising content according to that emotional state; means for analyzing the collected data to create a basic model of the virtual person's actions and voice; and means for adjusting the virtual person's actions and voice based on the emotion analysis data. This makes it possible to dynamically adjust the advertising content based on the emotions of each individual consumer and provide effective advertising in real time.
[0737] "A person's emotional state" refers to an individual's inner feelings and mood, as perceived from their facial expressions, voice, and other cues.
[0738] "Real-time recognition" refers to a process where information is processed as soon as it is generated, and results are obtained immediately.
[0739] "Dynamically adjusting ad content" means instantly changing the information presented as an ad according to the user's situation and emotions.
[0740] "Means of data collection" refers to any technical device or method for acquiring audio, image, or other sensory data.
[0741] A "foundation model" refers to a basic system or framework built to provide a foundation for a specific application.
[0742] "Emotional analysis data" refers to the results of data analysis used to identify a user's emotional state.
[0743] "Means for adjusting the behavior and voice of a virtual character" refers to a mechanism for changing the movements and speech patterns of a generated virtual character.
[0744] "Feedback data" refers to data collected from users' reactions and opinions to the information provided by the system.
[0745] A "media platform" is an online or offline technological infrastructure for distributing information and content.
[0746] A server plays a central role in implementing this invention. The server first collects the user's facial expressions and voice data through smart glasses or other devices. This data is analyzed in real time using emotion recognition software such as Google Cloud Vision or Amazon Rekognition to identify the user's emotional state. The analysis results are instantly reflected in the base model of the virtual person associated with each user.
[0747] On the server, the virtual character's actions and voice are adjusted based on analyzed emotion data. Using a generative AI model, the virtual character's behavior is designed to dynamically adjust the ad content according to the user's emotions. This virtual character is projected in real time onto smart glasses displays and other device screens, displaying ads in an interactive manner.
[0748] Furthermore, users can customize their virtual avatars through the interface, setting their appearance, voice, and behavior according to the company's marketing strategy. Feedback data is also collected and stored on the server for improvement.
[0749] For example, if a user wearing smart glasses is detected as being in a relaxed state, advertisements for calming relaxation products will be displayed with a bright and gentle voice. In this way, personalized advertising tailored to each user's emotional state is possible.
[0750] An example of a prompt for a generative AI model is, "Please think of the best advertisement to display when the user is relaxed."
[0751] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0752] Step 1:
[0753] The server collects video and audio data in real time from the user's smart glasses via the camera and microphone. This input data includes the user's facial expressions, voice tone, and tempo. This data forms the basis for analyzing the user's emotions.
[0754] Step 2:
[0755] The server inputs the collected video and audio data into emotion recognition software such as Google Cloud Vision or Amazon Rekognition to analyze the user's emotional state. This process uses machine learning algorithms to output emotion labels such as "happiness" and "stress" from the input data. This is then incorporated into the base model and used to adjust the virtual character.
[0756] Step 3:
[0757] The server uses a generative AI model to adjust the behavior and voice of the virtual character based on the analyzed emotional state. For example, if the user is relaxed, it generates a virtual character with a cheerful and calm voice and facial expression. This process also influences the selection of advertising content, leading to the output of optimal advertisements.
[0758] Step 4:
[0759] The device displays a customized virtual person and advertisement content on the smart glasses' display. In this step, the generated virtual person interacts with the user, creating a personalized advertising experience. Based on the output information from the server, the smart glasses play high-resolution video and audio.
[0760] Step 5:
[0761] Through the displayed advertising interface, users input settings for a virtual character and provide feedback on the ad content on their device. This feedback data is sent to the server and used to improve personalization in the future, contributing to overall system performance improvements.
[0762] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0763] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0764] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0765] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0766] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0767] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0768] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0769] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0770] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0771] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0772] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0773] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0774] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0775] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0776] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0777] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0778] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0779] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0780] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0781] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0782] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0783] The following is further disclosed regarding the embodiments described above.
[0784] (Claim 1)
[0785] A means of collecting data to generate virtual characters used in corporate advertising activities,
[0786] A means of analyzing collected data to create a basic model of the movements and voice of a virtual character,
[0787] A means of customizing virtual characters based on company requirements,
[0788] A means of generating the final virtual character based on customized settings,
[0789] A means of distributing the generated virtual character to multiple media platforms,
[0790] A means of improving the performance of a virtual character based on feedback data,
[0791] A system that includes this.
[0792] (Claim 2)
[0793] The system according to claim 1, which provides an interface for applying appearance, voice, or behavior to a virtual person in accordance with the specific requirements of a company.
[0794] (Claim 3)
[0795] The system according to claim 1, which automatically adjusts the variations of a virtual character by collecting and analyzing response data from an advertising campaign.
[0796] "Example 1"
[0797] (Claim 1)
[0798] A device for collecting information from a wide range of sources necessary for generating virtual characters,
[0799] A device for analyzing collected information and creating the basic structure of a virtual character's movements and voice,
[0800] A device for adjusting virtual characters to match a company's brand image,
[0801] A device for distributing customized virtual characters to various media outlets,
[0802] A device for continuously improving the performance of virtual characters based on feedback information from advertising campaigns,
[0803] A system that includes this.
[0804] (Claim 2)
[0805] The system according to claim 1, which provides an interface for reflecting the appearance, voice, and behavior of a virtual person in accordance with the specific requirements of a company.
[0806] (Claim 3)
[0807] The system according to claim 1, which automatically optimizes the variations of virtual characters by collecting and analyzing response information to advertising campaigns.
[0808] "Application Example 1"
[0809] (Claim 1)
[0810] A means of collecting information for generating a virtual character to be used in a medium,
[0811] A means of analyzing collected information to create a basic model of the movements and voice of a virtual character,
[0812] Means for adjusting virtual characters based on the requirements of the medium,
[0813] A means of generating the final virtual character based on adjusted settings,
[0814] A means of distributing the generated virtual person to multiple information distribution platforms,
[0815] A means of improving the performance of a virtual character based on feedback data,
[0816] A means of customizing virtual characters based on user interests and effectively conveying messages through videos and interactive simulations,
[0817] A system that includes this.
[0818] (Claim 2)
[0819] The system according to claim 1, which provides an interface for applying appearance, voice, or behavior to a virtual person in accordance with the specific requirements of a medium.
[0820] (Claim 3)
[0821] The system according to claim 1, comprising means for automatically adjusting variations of a virtual person and creating an impression by collecting and analyzing response data from an advertising campaign.
[0822] "Example 2 of combining an emotion engine"
[0823] (Claim 1)
[0824] A means for analyzing video and audio information obtained from the user to estimate the user's emotional state,
[0825] A means for automatically adjusting the actions and voice quality of a virtual character based on their estimated emotional state,
[0826] A means of generating a user emotion model and reflecting it in the representation of a virtual character,
[0827] A means of customizing virtual characters according to the characteristics of a company to generate virtual characters that are suitable for the company's advertising objectives,
[0828] A means of distributing generated virtual characters to various information media via information terminals,
[0829] Based on user feedback, a means to improve advertising efficiency by updating the behavior of virtual characters,
[0830] A system that includes this.
[0831] (Claim 2)
[0832] The system according to claim 1, comprising an operation screen for applying characteristics, voice, or behavior to a virtual person in accordance with the user's specific requests.
[0833] (Claim 3)
[0834] The system according to claim 1, which collects and analyzes user response information in advertising activities and enables dynamic adjustment of virtual characters.
[0835] "Application example 2 when combining with an emotional engine"
[0836] (Claim 1)
[0837] A means of collecting data to recognize a person's emotional state in real time and dynamically adjust advertising content according to that emotional state,
[0838] A means of analyzing collected data to create a basic model of the movements and voice of a virtual character,
[0839] A means of adjusting the behavior and voice of a virtual character based on emotion analysis data,
[0840] A means of customizing virtual characters based on company requirements,
[0841] A means of generating a final virtual character based on customized settings and displaying dynamic advertisements,
[0842] A means of distributing the generated virtual character to multiple media platforms,
[0843] A means to improve the performance of virtual characters based on feedback data and maximize the effectiveness of emotionally-driven advertising,
[0844] A system that includes this.
[0845] (Claim 2)
[0846] The system according to claim 1, which provides an interface for a virtual person that allows for adjustment of the appearance, voice, or actions corresponding to emotions.
[0847] (Claim 3)
[0848] The system according to claim 1, which automatically adjusts the variations of virtual characters by collecting and analyzing response data to an advertising campaign based on real-time human emotion data. [Explanation of symbols]
[0849] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of collecting data to generate virtual characters used in corporate advertising activities, A means of analyzing collected data to create a basic model of the movements and voice of a virtual character, A means of customizing virtual characters based on company requirements, A means of generating the final virtual character based on customized settings, A means of distributing the generated virtual character to multiple media platforms, A means of improving the performance of a virtual character based on feedback data, A system that includes this.
2. The system according to claim 1, which provides an interface for applying appearance, voice, or behavior to a virtual person in accordance with the specific requirements of a company.
3. The system according to claim 1, which automatically adjusts the variations of a virtual character by collecting and analyzing response data from an advertising campaign.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A