system

The system enhances audiovisual experiences by integrating haptic feedback, enabling users to feel the emotional impact of scenes through synchronized physical feedback, thus providing a richer and more immersive experience.

JP2026073367APending Publication Date: 2026-05-01SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Current audiovisual experiences primarily rely on sight and hearing, limiting the sense of presence and immersion, preventing users from fully enjoying the emotional impact and entertainment value of the content.

Method used

A system that integrates haptic feedback by generating representative sounds from audiovisual data, creating haptic signals, and transmitting composite data to a terminal device, allowing users to experience physical feedback synchronized with visual and auditory content.

Benefits of technology

Enhances the sensory experience by providing a richer, immersive entertainment experience that integrates sight, hearing, and touch, allowing users to feel the emotional impact of scenes through synchronized haptic feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026073367000001_ABST
    Figure 2026073367000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means of analyzing audiovisual data and using generation technology to generate representative sounds, A means for generating tactile signals from a generated representative sound, A data processing means for transmitting this composite data, including tactile signals, to a terminal device, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0005]

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

[0006] "Audiovisual data" is a general term for digital content that includes video and audio information.

[0007] "Analysis" is the act of processing input data in order to understand it structurally and semantically.

[0008] A "representative sound" is a key sound element associated with a particular video scene or audio content.

[0009] "Generative technology" refers to technologies that automatically create new data or signals using AI and algorithms.

[0010] A "tactile signal" is a command signal used to reproduce physical feedback through vibration or touch.

[0011] "Composite data" refers to a collection of multiple types of data, and in this context, it includes video, audio, and haptic data.

[0012] "Terminal device" refers to a playback device or computer device used to provide received data to the user.

[0013] A "physical device" refers to a device that receives tactile signals and reproduces physical vibrations or other sensations.

[0014] "User" refers to any person who uses this system to view and experience content. [Brief explanation of the drawing]

[0015] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Embodiments for Carrying Out the Invention

[0016] An example of an embodiment of the system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0019] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0020] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs, various parameters, and the like. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0023] [First Embodiment]

[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0036] The server receives audiovisual data such as movies, dramas, and documentaries. This data contains information about the images and sounds that viewers experience. The server uses generation technology to analyze this data and generate representative sounds corresponding to specific scenes. These representative sounds refer to sounds that strongly influence both visual and auditory perception, such as explosions in action scenes in movies or melodies in emotional scenes.

[0037] Based on the generated representative sound, the server creates a haptic signal. This haptic signal is intended to provide physical feedback, such as vibration, to the viewer and is transmitted to the user's terminal device. The server then constructs composite data, including the haptic signal, and transmits it to the terminal device through the distribution infrastructure.

[0038] The terminal receives composite data transmitted from the server, decomposes it, and obtains individual signals (video, audio, and haptic signals). The terminal plays back the video and audio signals, providing the user with a visual and auditory experience. The terminal also processes the haptic signals and generates vibrations via an onboard physical device. This feedback is provided in real time, synchronized with the movie scenes.

[0039] Users select video content using their terminal device and receive various media information visually and aurally. Haptic feedback allows users to experience a sense of reality and immersion that cannot be obtained through video and audio alone. For example, when watching an action movie, the device vibrates during explosion scenes, providing a more immersive experience.

[0040] In this way, the present invention aims to realize a new media experience that integrates the three senses of sight, hearing, and touch, and to provide users with unprecedented immersive entertainment.

[0041] The following describes the processing flow.

[0042] Step 1:

[0043] The server receives audiovisual data from video streaming service providers. This data includes video and audio information, which is then analyzed and processed.

[0044] Step 2:

[0045] The server receives audiovisual data and analyzes it using generation technology. Here, it identifies important scenes and generates representative sounds corresponding to them.

[0046] Step 3:

[0047] The server creates haptic signals based on the generated representative sound. This includes setting specific vibration patterns. For example, a strong, rapid vibration might be set for an explosion sound.

[0048] Step 4:

[0049] The server integrates video, audio, and haptic signals to create a composite data package. This data includes visual, auditory, and tactile information.

[0050] Step 5:

[0051] The server sends the generated composite data package to the terminal device. This transmission is performed using the network infrastructure and is delivered in real-time or on-demand.

[0052] Step 6:

[0053] The terminal receives the composite data sent from the server. The data arrives via the communication module and is ready for use.

[0054] Step 7:

[0055] The terminal analyzes the received composite data and separates it into video, audio, and haptic signals. These signals are then distributed to their respective playback systems.

[0056] Step 8:

[0057] The terminal sends video and audio signals to corresponding output devices (displays and speakers) and displays and plays them in a format that the user can view.

[0058] Step 9:

[0059] The device sends haptic signals to a vibration motor, generating physical vibrations in real time. This allows the user to receive haptic feedback.

[0060] Step 10:

[0061] Users enjoy audiovisual information provided through their devices and gain immersion through haptic feedback. Haptic feedback has the effect of amplifying the impact and emotion of the scenes.

[0062] (Example 1)

[0063] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0064] Current audiovisual experiences are limited to sight and hearing, which presents a challenge in that users cannot deeply feel a sense of presence or immersion in the content. This challenge can prevent viewers from fully enjoying the emotional impact and entertainment value of the content.

[0065] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0066] In this invention, the server includes means for using generation technology to analyze audiovisual information and generate characteristic sounds, means for creating vibration signals based on the generated sounds, and information processing means for transmitting composite information including vibration signals to an information processing device. This enables the user to have a new media experience that integrates the three senses of sight, hearing, and touch.

[0067] "Audiovisual information" is a general term for video and audio data acquired through sight and hearing.

[0068] "Distinctive sound" refers to an audio signal generated to emphasize a specific scene within audiovisual information.

[0069] "Generative technology" refers to the technology of analyzing data and creating new acoustic signals.

[0070] A "vibration signal" is a signal used to generate physical vibrations.

[0071] "Complex information" refers to a dataset that integrates video signals, audio signals, vibration signals, and other data.

[0072] An "information processing device" is a device that decomposes received complex information and distributes each signal to the appropriate output device.

[0073] A "physical device" is a device that receives vibration signals and generates actual vibrations.

[0074] "Sensory experience" refers to a new experience that a user gains through sight, hearing, and touch.

[0075] In implementing this invention, the server plays the role of receiving audiovisual information. Specifically, it acquires visual and auditory data from content such as movies, dramas, and documentaries. The server analyzes this data using a generative AI model and generates characteristic sounds associated with a specific scene. An example of a prompt sentence to be input to the AI ​​model is, "Generate representative sounds from an action scene in a movie."

[0076] Based on the generated sound, the server creates a vibration signal. For example, if there is an explosion scene in a movie, a strong vibration signal will be generated. This vibration signal is combined with the video and audio signals as composite information and transmitted from the server to the terminal, which is an information processing device.

[0077] The device decomposes the received composite information and outputs video signals to the display and audio signals to the speaker, providing the user with an audiovisual experience. Furthermore, vibration signals are reproduced as actual vibrations by a physical device built into the device. In this way, the user can obtain a sensory experience that utilizes not only sight and hearing, but also touch. This haptic feedback enhances the sense of presence and immersion in the content, providing a richer entertainment experience.

[0078] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0079] Step 1:

[0080] The server receives audiovisual information. This data includes video and audio data from movies, dramas, and documentaries. This data is used as input for analysis by a generative AI model.

[0081] Step 2:

[0082] The server analyzes the input audiovisual information using a generative AI model. The AI ​​model receives the prompt "Generate representative sounds from an action scene in a movie," and based on this instruction, identifies a specific scene from the audiovisual data and generates characteristic sounds. The output is characteristic sound data for that scene.

[0083] Step 3:

[0084] The server receives the generated acoustic data as input and creates a vibration signal based on it. For example, if the generated sound is an explosion, a strong vibration signal is created. This vibration signal is output as partial data for later composite information construction.

[0085] Step 4:

[0086] The server integrates video, audio, and vibration signals to form composite information. This composite information becomes output data to be transmitted to the terminal, which is an information processing device.

[0087] Step 5:

[0088] The terminal receives composite information transmitted from the server. This information includes video signals, audio signals, and vibration signals, which are then separated and used as input data to be appropriately distributed to each output device.

[0089] Step 6:

[0090] The device outputs the decomposed video signal to the display and the audio signal to the speaker. This provides the user with an audiovisual experience. The output here is the actual content that the user experiences visually and aurally.

[0091] Step 7:

[0092] The device receives vibration signals and reproduces actual vibrations using its built-in physical components. The device then delivers physical vibrations as output, allowing the user to experience realistic feedback through touch. This process enables the user to obtain an immersive and realistic media experience.

[0093] (Application Example 1)

[0094] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0095] Modern audiovisual content, which relies solely on video and audio for the experience, limits interaction with the real world and a sense of presence. There is a need for methods to provide users with more immersive and sensory-rich experiences.

[0096] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0097] In this invention, the server includes means for analyzing audiovisual information and using a generation method to generate characteristic sounds, means for generating tactile signals from the generated characteristic sounds, and information processing means for transmitting composite information including these tactile signals to an information terminal. This makes it possible to improve the user's sensory experience by providing tactile feedback.

[0098] "Audiovisual information" refers to digital data related to sight and hearing, including images and sounds.

[0099] A "characteristic sound" is an acoustic signal that represents a specific scene within audiovisual information and is generated to enhance the user experience.

[0100] A "generation method" is a computational procedure or algorithm used to achieve a specific purpose, and is a technique used to generate characteristic sounds from audiovisual information.

[0101] "Haptic signals" are signals that elicit physical feedback, such as vibration or pressure, and provide users with tactile interaction.

[0102] "Composite information" refers to a collection of information including images, sounds, and haptic signals, which is transmitted to an information terminal to provide users with a multi-sensory experience.

[0103] An "information terminal" is an electronic device equipped with the function of displaying and playing content, which processes the received composite information to provide an experience to the user.

[0104] The system implementing this invention mainly consists of a server and an information terminal. The server analyzes audiovisual information using a generative AI model and generates characteristic sounds that represent a specific scene. Based on these characteristic sounds, the server generates tactile signals and transmits them to the information terminal as composite information.

[0105] The information terminal decomposes the received composite information and plays back video and audio to provide the user with an audiovisual experience. Furthermore, it realizes haptic signals through a physical device, providing the user with real-time haptic feedback. This feedback allows the user to experience a sense of presence that cannot be obtained from video and audio alone.

[0106] Implementing this system requires information terminals such as smartphones or tablets equipped with vibration motors. The server software uses a generative AI model to analyze audiovisual information and generate characteristic sounds. In particular, APIs for generating tactile signals, such as OpenHaptics, are utilized.

[0107] For example, during an action scene in a movie, the smartphone might vibrate strongly, allowing the user to feel the tension of that scene. This kind of feedback provides a new entertainment experience that integrates not only sight and hearing but also touch.

[0108] An example of a prompt message sent to a generative AI model is: "The next scene is an emotional climax. Please generate subtle haptic feedback that will evoke emotions in the user."

[0109] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0110] Step 1:

[0111] The server receives audiovisual information. It takes video and audio data from movies and dramas as input and analyzes this data using a generative AI model. As a result of the analysis, it outputs characteristic sounds that represent a specific scene.

[0112] Step 2:

[0113] The server generates haptic signals based on characteristic sounds. It takes characteristic sounds as input and outputs the generated haptic signals. Here, vibration patterns are designed using APIs such as OpenHaptics to generate haptic signals suitable for feedback.

[0114] Step 3:

[0115] The server combines the generated haptic signals with video and audio data to form composite information. This composite information is the data output and is prepared for transmission to information terminals.

[0116] Step 4:

[0117] The information terminal decomposes the complex information received from the server. It separates the complex information received as input into video signals, audio signals, and haptic signals, and processes each of these individually.

[0118] Step 5:

[0119] The information terminal plays video and audio signals, providing the user with an audiovisual experience. In this step, the media player on the information terminal operates, providing the user with audiovisual content from video and audio.

[0120] Step 6:

[0121] The information terminal executes haptic signals via a physical device. Here, haptic signals are input to a physical device such as a vibration motor, and a specified vibration feedback is provided to the user as output.

[0122] Step 7:

[0123] Users enjoy a tripartite sensory experience—audiovisual and tactile—provided by the information terminal. In this way, they experience immersive entertainment with haptic feedback optimized for each scene in movies and dramas.

[0124] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0125] The server receives audiovisual data from video content creators, which includes video and audio information. First, the server analyzes this audiovisual data using generation technology. It identifies important scenes and events and generates corresponding representative sounds. By creating haptic signals based on the generated representative sounds, it prepares to provide users with haptic feedback that is linked to their audiovisual experience.

[0126] The server further utilizes an emotion engine to analyze the user's emotions in real time. The emotion engine uses data such as the user's facial expressions, heart rate, and speech to understand their current emotional state. Based on this emotional information, the server dynamically adjusts haptic signals to create haptic feedback that is optimal for the user's current emotions.

[0127] The generated haptic signals are packaged together with visual and auditory signals as composite data and transmitted to a terminal device. The terminal receives this composite data and plays the video and audio signals, while also processing the haptic signals and implementing them with a vibration motor. This allows the user to experience a deeper sense of immersion in the content they are viewing.

[0128] For example, when a user is watching an emotionally moving film, the emotion engine detects the user's tears or smiles and adjusts the haptic signals to a milder, more pleasant vibration. Similarly, during action scenes, it senses the user's excitement level and provides strong, vivid vibration feedback. This real-time emotional feedback allows users to enjoy a more interactive and personalized entertainment experience.

[0129] This technology expands the possibilities of unprecedented viewing experiences by integrating visual, auditory, and tactile senses, as well as providing adaptive feedback tailored to the user's emotional state. Seamless communication between the server and the device, along with the recognition and application of emotions, allows users to obtain individually optimized content experiences.

[0130] The following describes the processing flow.

[0131] Step 1:

[0132] The server receives audiovisual data, including video and audio, from video streaming providers. The received data is prepared for analysis.

[0133] Step 2:

[0134] The server uses generation technology to analyze audiovisual data and generate representative sounds corresponding to specific scenes. This analysis identifies key events in the video and generates acoustic information based on them.

[0135] Step 3:

[0136] The server creates a haptic signal based on the generated representative sound. This haptic signal includes an appropriate vibration pattern corresponding to the scene of the content being viewed.

[0137] Step 4:

[0138] The server uses an emotion engine to analyze the user's emotions in real time. Based on sensor data received from the user (e.g., camera footage and heart rate), it analyzes the user's emotional state and creates an emotion profile.

[0139] Step 5:

[0140] The server adjusts haptic signals based on the user's emotional profile. This dynamically changes the intensity and pattern of vibrations according to the user's emotional state to provide optimal feedback.

[0141] Step 6:

[0142] The server creates a composite data package containing video, audio, and tuned haptic signals. This package is then prepared in the appropriate format for transmission to the terminal device.

[0143] Step 7:

[0144] The server transmits the composite data package to the terminal device. The transmission takes place over the network, ensuring that the data is transmitted correctly and without delay.

[0145] Step 8:

[0146] The terminal receives the composite data transmitted from the server. After receiving, the data is analyzed for processing and separated into its respective signal formats.

[0147] Step 9:

[0148] The terminal transmits separated video and audio signals to the playback device, providing the user with visual and auditory feedback.

[0149] Step 10:

[0150] The device transmits haptic signals to a vibration motor, generating vibrations that allow the user to experience haptic feedback. This feedback is provided in real time.

[0151] Step 11:

[0152] Users enjoy the viewing experience through the video, audio, and haptic feedback provided by the device. Emotion-based haptic adjustments allow for a more immersive experience.

[0153] (Example 2)

[0154] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0155] Traditional audiovisual content has relied on visual and auditory experiences, but it lacks mechanisms for users to become more emotionally immersed in the content and engage in interactive experiences. In particular, there is a need to enhance the sensory experience by providing tactile feedback in addition to visual and auditory information. Furthermore, a challenge lies in achieving more personalized responses by dynamically adjusting tactile signals according to the user's emotional state.

[0156] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0157] In this invention, the server includes means for using a generation method to analyze audiovisual information, identify specific scenes, and generate acoustic information; means for creating haptic signals from the generated acoustic information; and analysis device means for analyzing the user's emotional state and dynamically adjusting the haptic signals based on that state. This enables the user to obtain an immersive experience that integrates visual, auditory, and tactile senses with the content they are viewing, and in addition, enables personalized haptic feedback that responds to the user's real-time emotional state.

[0158] "Audiovisual information" is a general term for information that includes visual data related to vision and audio data related to hearing.

[0159] A "generation method" is a technical means for identifying specific scenes from audiovisual information and creating audio information tailored to a particular purpose.

[0160] "Acoustic information" refers to sound information, including music and sound effects appropriate to a scene, generated through the analysis of audiovisual information.

[0161] "Haptic signals" are signals that represent physical stimuli such as vibration and pressure, designed to provide haptic feedback to the user.

[0162] An "analysis device" refers to a device or software that has the function of collecting and analyzing real-time data from users and evaluating their emotional state.

[0163] "Combined information" refers to a dataset that integrates visual, auditory, and tactile information, which is transmitted to the user's terminal device to provide a holistic experience.

[0164] A "receiving device" refers to a device that receives composite information transmitted from a server, and is used to provide users with audiovisual experiences and haptic feedback.

[0165] An "output device" refers to a device that transmits physical haptic feedback to the user based on the received haptic signals.

[0166] "Perceptual experience" refers to an interactive experience provided to the user through an integrated sensory experience encompassing sight, hearing, and touch.

[0167] This invention relates to a system for enhancing the user experience based on audiovisual information. This system consists of a server, a terminal, and a user, and each component works in conjunction to provide an unprecedented multi-sensory experience.

[0168] The server receives audiovisual information and analyzes that data. In particular, it uses generation techniques to identify specific scenes from the audiovisual information and generate appropriate audio information. The server can use image analysis software and audio analysis tools for the processing required for analysis. For example, it can use OpenCV, an open-source image processing library, and libraries for extracting audio features.

[0169] The server then creates haptic signals from the generated acoustic information. AI technology is used to design haptic feedback that is synchronized with the acoustic information. Since the haptic signals are physically fed back to the user as vibrations and pressure, their design is crucial.

[0170] Furthermore, the server analyzes the user's emotional state. To do this, it uses an emotion analysis device to analyze the user's real-time data (e.g., camera footage and audio data). This allows the server to dynamically adjust haptic signals in response to changes in the user's emotions. This dynamic adjustment makes it possible to provide the user with a personalized haptic experience.

[0171] The server transmits visual, auditory, and dynamically adjusted haptic signals as a composite information to the terminal. This composite information is transmitted quickly and accurately using real-time communication technology. Low-latency methods such as WebSocket are suitable as the communication protocol.

[0172] The device decomposes the received composite information and reproduces the visual and auditory information on the user's screen or speaker. Meanwhile, haptic signals are used to drive a vibration motor, providing haptic feedback. The vibration output device included in the device delivers accurate haptic feedback to the user based on a pre-designed vibration pattern.

[0173] Users can enjoy an integrated visual, auditory, and tactile experience delivered through the device. This system provides a deeper sense of immersion and a more personalized interactive experience by offering adaptive feedback that responds to the user's emotions.

[0174] As a concrete example of a prompt, one could input the following into the generating AI model: "Explain how to dynamically adjust and display haptic feedback according to the user's emotions." By designing the operation according to this example, it becomes possible to realize the new user experience that the system aims for.

[0175] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0176] Step 1:

[0177] The server receives audiovisual data from video content creators. This data includes video and audio data. The received audiovisual data is converted to a different format by the server and prepared for analysis.

[0178] Step 2:

[0179] The server analyzes audiovisual data to identify specific scenes. This analysis uses image analysis software to analyze video data frame by frame and select specific scenes. It also uses audio analysis tools to extract features from audio data. The input to this process is video data and audio data, and the output is identified scene information and audio feature data.

[0180] Step 3:

[0181] The server generates acoustic information using a generation method based on the identified scene. A generation AI model is utilized to create music and sound effects that are optimal for the selected scene. The input is identified scene information and audio feature data, and the output is the generated acoustic information.

[0182] Step 4:

[0183] The server creates haptic signals based on the generated acoustic information. The haptic signals are data representations of vibration patterns designed to match the acoustic information. In this step, the generated acoustic information is used as input, and haptic signal data is obtained as output.

[0184] Step 5:

[0185] The server collects real-time emotional data from users and performs emotional analysis. This includes analyzing camera footage and audio input to evaluate the user's emotional state. The input is real-time data from the user, and the output is emotional state information obtained through analysis.

[0186] Step 6:

[0187] The server dynamically adjusts haptic signals based on the user's emotional state information. The vibration intensity and pattern of the haptic signals are modified according to the analysis results. The input consists of emotional state information and haptic signal data, while the output is the adjusted haptic signal.

[0188] Step 7:

[0189] The server packages the tuned haptic signals and audiovisual information as composite data and transmits it to the terminal. A low-latency communication protocol is used to transmit the composite data. The input is the tuned haptic signals and audiovisual information, and the output is the data to be transmitted to the terminal.

[0190] Step 8:

[0191] The terminal decomposes the received composite information and reproduces the visual and auditory information for the user. Simultaneously, it transmits haptic signals to a vibration device, providing the user with physical feedback. The input is data transmitted from the server, and the output is video, audio, and haptic feedback.

[0192] (Application Example 2)

[0193] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0194] Modern content delivery services demand immersive experiences that integrate not only audiovisual but also tactile senses. However, current technology makes it difficult to provide dynamic haptic feedback in real time that responds to the user's emotions. Therefore, there are technical challenges to further enhance the viewer's sensory experience.

[0195] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0196] In this invention, the server includes means for using generation technology to analyze audiovisual signals and generate representative acoustic information, means for creating tactile information based on the generated acoustic information, and data processing means for dynamically adjusting the tactile information based on the user's emotions and transmitting the composite information to a terminal device. This makes it possible to provide the user with real-time tactile feedback synchronized with audiovisual data and optimize the perceptual experience.

[0197] "Audiovisual signals" are data that includes image and sound information, and are fundamental elements for users to experience content through sight and hearing.

[0198] "Acoustic information" is data that represents the characteristics of sound, generated by analyzing audiovisual signals, and extracts representative sounds and important sound changes.

[0199] "Haptic information" refers to signals generated to provide feedback that directly appeals to the user's senses, reproducing physical sensations such as vibrations.

[0200] "User emotions" refer to information that indicates the viewer's current feelings and mental state, and are analyzed in real time from facial expressions and biosignals.

[0201] "Composite information" refers to a data package that integrates audiovisual signals, acoustic information, and haptic information, and is transmitted to terminal devices to provide users with an immersive experience.

[0202] "Terminal device" refers to a device used by a user to experience content, and includes smartphones, head-mounted displays, and other similar devices.

[0203] "Data processing" refers to a series of computational processes that analyze, integrate, and transmit multiple pieces of information, and is a central function of a system for optimizing the user experience.

[0204] The system of the present invention aims to provide users with an immersive experience that integrates sight, hearing, and touch. This system mainly consists of two main components: a server and a terminal device. The server receives audiovisual signals, analyzes important scenes using generation technology, and generates representative acoustic information. Subsequently, it creates tactile information based on this acoustic information and dynamically adjusts the tactile information using an emotion engine that analyzes the user's emotions in real time. This adjusted tactile information is packaged together with the audiovisual signals as composite information and transmitted to the terminal device.

[0205] The terminal device receives this composite information, plays back the audio information, and simultaneously materializes the haptic information using a physical device. This allows users to enhance their content experience not only through sight and hearing, but also through touch. Specific hardware includes smartphones and head-mounted displays, while software such as EmotionEngine and HapticFeedback is used.

[0206] As a concrete example, when a user is watching an emotionally moving film, the server analyzes the viewer's facial expressions, captures subtle emotional changes, and adjusts the haptic information accordingly. Soft vibrations are provided during emotional scenes, while stronger vibrations are generated during action scenes, resulting in an interactive experience.

[0207] Another example of providing prompts to a generative AI model is, "Design a system that analyzes user emotions in real time and provides dynamic haptic feedback according to the video scene." Based on this prompt, the AI ​​model generates an appropriate response, enabling the design of enhanced systems.

[0208] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0209] Step 1:

[0210] The server receives audiovisual signals from content creators. These signals include the video and audio that the user will view. The server uses generation technology to analyze these audiovisual signals and identify important scenes. This analysis allows the server to extract video features and acoustic information and generate representative acoustic information corresponding to the important scenes.

[0211] Step 2:

[0212] The server generates haptic information based on the generated acoustic information. This haptic information is designed to provide the user with a sensory experience in the form of vibration and pressure. This process analyzes the intensity and changes in the acoustic information and designs corresponding vibration patterns. As a result, haptic information that the user actually feels is constructed.

[0213] Step 3:

[0214] The server uses an emotion engine to analyze the user's emotions in real time. It receives the user's facial expression data and biosignals as input and evaluates the current emotional state based on that data. The results of the emotion analysis are used to dynamically adjust haptic information. This makes it possible to design optimal haptic feedback that matches the user's emotions.

[0215] Step 4:

[0216] The server integrates haptic information with audiovisual signals and packages it as composite information. This process seamlessly integrates haptic, acoustic, and visual information to generate a dataset that provides the user with a more complete sensory experience. The generated composite information is then ready to be transmitted to the terminal device.

[0217] Step 5:

[0218] The terminal receives composite information transmitted from the server. It decomposes the composite information, plays the acoustic information through an audio output device, and displays the visual information on the display. It also analyzes tactile information and converts it into signals to control a vibration motor, providing the user with physical tactile feedback.

[0219] Step 6:

[0220] Users enjoy a multi-sensory experience through their devices. By experiencing content that integrates sight, hearing, and touch, they can achieve an unprecedented level of immersion. Through this process, users can feel as if they are actually inside the video.

[0221] This series of processes allows for the dynamic adjustment of haptic information based on the user's emotions, while utilizing audiovisual signals, to provide a personalized entertainment experience.

[0222] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0223] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0224] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0225] [Second Embodiment]

[0226] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0227] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0228] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0229] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0230] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0231] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0232] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0233] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0234] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0235] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0236] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0237] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0238] The server receives audiovisual data such as movies, dramas, and documentaries. This data contains information about the images and sounds that viewers experience. The server uses generation technology to analyze this data and generate representative sounds corresponding to specific scenes. These representative sounds refer to sounds that strongly influence both visual and auditory perception, such as explosions in action scenes in movies or melodies in emotional scenes.

[0239] Based on the generated representative sound, the server creates a haptic signal. This haptic signal is intended to provide physical feedback, such as vibration, to the viewer and is transmitted to the user's terminal device. The server then constructs composite data, including the haptic signal, and transmits it to the terminal device through the distribution infrastructure.

[0240] The terminal receives composite data transmitted from the server, decomposes it, and obtains individual signals (video, audio, and haptic signals). The terminal plays back the video and audio signals, providing the user with a visual and auditory experience. The terminal also processes the haptic signals and generates vibrations via an onboard physical device. This feedback is provided in real time, synchronized with the movie scenes.

[0241] Users select video content using their terminal device and receive various media information visually and aurally. Haptic feedback allows users to experience a sense of reality and immersion that cannot be obtained through video and audio alone. For example, when watching an action movie, the device vibrates during explosion scenes, providing a more immersive experience.

[0242] In this way, the present invention aims to realize a new media experience that integrates the three senses of sight, hearing, and touch, and to provide users with unprecedented immersive entertainment.

[0243] The following describes the processing flow.

[0244] Step 1:

[0245] The server receives audiovisual data from video streaming service providers. This data includes video and audio information, which is then analyzed and processed.

[0246] Step 2:

[0247] The server receives audiovisual data and analyzes it using generation technology. Here, it identifies important scenes and generates representative sounds corresponding to them.

[0248] Step 3:

[0249] The server creates haptic signals based on the generated representative sound. This includes setting specific vibration patterns. For example, a strong, rapid vibration might be set for an explosion sound.

[0250] Step 4:

[0251] The server integrates video, audio, and haptic signals to create a composite data package. This data includes visual, auditory, and tactile information.

[0252] Step 5:

[0253] The server sends the generated composite data package to the terminal device. This transmission is performed using the network infrastructure and is delivered in real-time or on-demand.

[0254] Step 6:

[0255] The terminal receives the composite data sent from the server. The data arrives via the communication module and is ready for use.

[0256] Step 7:

[0257] The terminal analyzes the received composite data and separates it into video, audio, and haptic signals. These signals are then distributed to their respective playback systems.

[0258] Step 8:

[0259] The terminal sends video and audio signals to corresponding output devices (displays and speakers) and displays and plays them in a format that the user can view.

[0260] Step 9:

[0261] The device sends haptic signals to a vibration motor, generating physical vibrations in real time. This allows the user to receive haptic feedback.

[0262] Step 10:

[0263] Users enjoy audiovisual information provided through their devices and gain immersion through haptic feedback. Haptic feedback has the effect of amplifying the impact and emotion of the scenes.

[0264] (Example 1)

[0265] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0266] Current audiovisual experiences are limited to sight and hearing, which presents a challenge in that users cannot deeply feel a sense of presence or immersion in the content. This challenge can prevent viewers from fully enjoying the emotional impact and entertainment value of the content.

[0267] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0268] In this invention, the server includes means for using generation technology to analyze audiovisual information and generate characteristic sounds, means for creating vibration signals based on the generated sounds, and information processing means for transmitting composite information including vibration signals to an information processing device. This enables the user to have a new media experience that integrates the three senses of sight, hearing, and touch.

[0269] "Audiovisual information" is a general term for video and audio data acquired through sight and hearing.

[0270] "Distinctive sound" refers to an audio signal generated to emphasize a specific scene within audiovisual information.

[0271] "Generative technology" refers to the technology of analyzing data and creating new acoustic signals.

[0272] A "vibration signal" is a signal used to generate physical vibrations.

[0273] "Complex information" refers to a dataset that integrates video signals, audio signals, vibration signals, and other data.

[0274] An "information processing device" is a device that decomposes received complex information and distributes each signal to the appropriate output device.

[0275] A "physical device" is a device that receives vibration signals and generates actual vibrations.

[0276] "Sensory experience" refers to a new experience that a user gains through sight, hearing, and touch.

[0277] In implementing this invention, the server plays the role of receiving audiovisual information. Specifically, it acquires visual and auditory data from content such as movies, dramas, and documentaries. The server analyzes this data using a generative AI model and generates characteristic sounds associated with a specific scene. An example of a prompt sentence to be input to the AI ​​model is, "Generate representative sounds from an action scene in a movie."

[0278] Based on the generated sound, the server creates a vibration signal. For example, if there is an explosion scene in a movie, a strong vibration signal will be generated. This vibration signal is combined with the video and audio signals as composite information and transmitted from the server to the terminal, which is an information processing device.

[0279] The terminal decomposes the received composite information and outputs the video signal to the display and the audio signal to the speaker, providing the user with a visual and auditory experience. Also, the vibration signal is reproduced as an actual vibration by a physical device built into the terminal. Thus, the user can obtain a sensory experience that utilizes the tactile sense in addition to vision and hearing. This tactile feedback enhances the sense of presence and immersion in the content, providing a richer entertainment experience.

[0280] The flow of the specific process in Example 1 will be described using FIG. 11.

[0281] Step 1:

[0282] The server receives visual and auditory information. The data to be received includes video and audio data of movies, dramas, and documentaries. These data are used as inputs for analysis by the generation AI model.

[0283] Step 2:

[0284] The server analyzes the input visual and auditory information using the generation AI model. The AI model receives, as a prompt sentence, "Please generate representative sounds in the action scenes of the movie" and, based on that instruction, identifies specific scenes from the visual data and generates characteristic sounds. The output is characteristic acoustic data for the scenes.

[0285] Step 3:

[0286] The server receives the generated acoustic data as input and creates a vibration signal based on it. For example, if the generated sound is an explosion sound, a strong vibration signal is created. This vibration signal is output as partial data for later composite information composition.

[0287] Step 4:

[0288] The server integrates video, audio, and vibration signals to form composite information. This composite information becomes output data to be transmitted to the terminal, which is an information processing device.

[0289] Step 5:

[0290] The terminal receives composite information transmitted from the server. This information includes video signals, audio signals, and vibration signals, which are then separated and used as input data to be appropriately distributed to each output device.

[0291] Step 6:

[0292] The device outputs the decomposed video signal to the display and the audio signal to the speaker. This provides the user with an audiovisual experience. The output here is the actual content that the user experiences visually and aurally.

[0293] Step 7:

[0294] The device receives vibration signals and reproduces actual vibrations using its built-in physical components. The device then delivers physical vibrations as output, allowing the user to experience realistic feedback through touch. This process enables the user to obtain an immersive and realistic media experience.

[0295] (Application Example 1)

[0296] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0297] Modern audiovisual content, which relies solely on video and audio for the experience, limits interaction with the real world and a sense of presence. There is a need for methods to provide users with more immersive and sensory-rich experiences.

[0298] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0299] In this invention, the server includes means for analyzing audiovisual information and using a generation method to generate characteristic sounds, means for generating tactile signals from the generated characteristic sounds, and information processing means for transmitting composite information including these tactile signals to an information terminal. This makes it possible to improve the user's sensory experience by providing tactile feedback.

[0300] "Audiovisual information" refers to digital data related to sight and hearing, including images and sounds.

[0301] A "characteristic sound" is an acoustic signal that represents a specific scene within audiovisual information and is generated to enhance the user experience.

[0302] A "generation method" is a computational procedure or algorithm used to achieve a specific purpose, and is a technique used to generate characteristic sounds from audiovisual information.

[0303] "Haptic signals" are signals that elicit physical feedback, such as vibration or pressure, and provide users with tactile interaction.

[0304] "Composite information" refers to a collection of information including images, sounds, and haptic signals, which is transmitted to an information terminal to provide users with a multi-sensory experience.

[0305] An "information terminal" is an electronic device equipped with the function of displaying and playing content, which processes the received composite information to provide an experience to the user.

[0306] The system implementing this invention mainly consists of a server and an information terminal. The server analyzes audiovisual information using a generative AI model and generates characteristic sounds that represent a specific scene. Based on these characteristic sounds, the server generates tactile signals and transmits them to the information terminal as composite information.

[0307] The information terminal decomposes the received composite information, plays back the video and audio, and provides the user with an audiovisual experience. Furthermore, it realizes tactile signals through a physical device and provides the user with tactile feedback in real time. With this feedback, the user can obtain a sense of presence that cannot be obtained only from video and audio.

[0308] To implement this system, an information terminal such as a smartphone or tablet equipped with a vibration motor is required. For the server software, a generation AI model for analyzing audiovisual information and generating characteristic sounds is used. In particular, an API for generating tactile signals such as OpenHaptics is utilized.

[0309] As a specific example, when watching a movie, the smartphone vibrates strongly during an action scene, enabling the user to experience the tension of that scene. Such feedback provides a new entertainment experience that integrates not only vision and hearing but also touch.

[0310] An example of the prompt text sent to the generation AI model is "The next scene is an emotional climax. Please generate delicate tactile feedback that evokes emotions in the user."

[0311] The flow of the specific process in Application Example 1 will be described using FIG. 12.

[0312] Step 1:

[0313] The server receives audiovisual information. It receives video and audio data of a movie or drama as input and analyzes this data using the generation AI model. As a result of the analysis, it outputs characteristic sounds representing a specific scene.

[0314] Step 2:

[0315] The server generates haptic signals based on characteristic sounds. It takes characteristic sounds as input and outputs the generated haptic signals. Here, vibration patterns are designed using APIs such as OpenHaptics to generate haptic signals suitable for feedback.

[0316] Step 3:

[0317] The server combines the generated haptic signals with video and audio data to form composite information. This composite information is the data output and is prepared for transmission to information terminals.

[0318] Step 4:

[0319] The information terminal decomposes the complex information received from the server. It separates the complex information received as input into video signals, audio signals, and haptic signals, and processes each of these individually.

[0320] Step 5:

[0321] The information terminal plays video and audio signals, providing the user with an audiovisual experience. In this step, the media player on the information terminal operates, providing the user with audiovisual content from video and audio.

[0322] Step 6:

[0323] The information terminal executes haptic signals via a physical device. Here, haptic signals are input to a physical device such as a vibration motor, and a specified vibration feedback is provided to the user as output.

[0324] Step 7:

[0325] Users enjoy a tripartite sensory experience—audiovisual and tactile—provided by the information terminal. In this way, they experience immersive entertainment with haptic feedback optimized for each scene in movies and dramas.

[0326] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0327] The server receives audiovisual data from video content creators, which includes video and audio information. First, the server analyzes this audiovisual data using generation technology. It identifies important scenes and events and generates corresponding representative sounds. By creating haptic signals based on the generated representative sounds, it prepares to provide users with haptic feedback that is linked to their audiovisual experience.

[0328] The server further utilizes an emotion engine to analyze the user's emotions in real time. The emotion engine uses data such as the user's facial expressions, heart rate, and speech to understand their current emotional state. Based on this emotional information, the server dynamically adjusts haptic signals to create haptic feedback that is optimal for the user's current emotions.

[0329] The generated haptic signals are packaged together with visual and auditory signals as composite data and transmitted to a terminal device. The terminal receives this composite data and plays the video and audio signals, while also processing the haptic signals and implementing them with a vibration motor. This allows the user to experience a deeper sense of immersion in the content they are viewing.

[0330] For example, when a user is watching an emotionally moving film, the emotion engine detects the user's tears or smiles and adjusts the haptic signals to a milder, more pleasant vibration. Similarly, during action scenes, it senses the user's excitement level and provides strong, vivid vibration feedback. This real-time emotional feedback allows users to enjoy a more interactive and personalized entertainment experience.

[0331] This technology expands the possibilities of unprecedented viewing experiences by integrating visual, auditory, and tactile senses, as well as providing adaptive feedback tailored to the user's emotional state. Seamless communication between the server and the device, along with the recognition and application of emotions, allows users to obtain individually optimized content experiences.

[0332] The following describes the processing flow.

[0333] Step 1:

[0334] The server receives audiovisual data, including video and audio, from video streaming providers. The received data is prepared for analysis.

[0335] Step 2:

[0336] The server uses generation technology to analyze audiovisual data and generate representative sounds corresponding to specific scenes. This analysis identifies key events in the video and generates acoustic information based on them.

[0337] Step 3:

[0338] The server creates a haptic signal based on the generated representative sound. This haptic signal includes an appropriate vibration pattern corresponding to the scene of the content being viewed.

[0339] Step 4:

[0340] The server uses an emotion engine to analyze the user's emotions in real time. Based on sensor data received from the user (e.g., camera footage and heart rate), it analyzes the user's emotional state and creates an emotion profile.

[0341] Step 5:

[0342] The server adjusts haptic signals based on the user's emotional profile. This dynamically changes the intensity and pattern of vibrations according to the user's emotional state to provide optimal feedback.

[0343] Step 6:

[0344] The server creates a composite data package containing video, audio, and tuned haptic signals. This package is then prepared in the appropriate format for transmission to the terminal device.

[0345] Step 7:

[0346] The server transmits the composite data package to the terminal device. The transmission takes place over the network, ensuring that the data is transmitted correctly and without delay.

[0347] Step 8:

[0348] The terminal receives the composite data transmitted from the server. After receiving, the data is analyzed for processing and separated into its respective signal formats.

[0349] Step 9:

[0350] The terminal transmits separated video and audio signals to the playback device, providing the user with visual and auditory feedback.

[0351] Step 10:

[0352] The device transmits haptic signals to a vibration motor, generating vibrations that allow the user to experience haptic feedback. This feedback is provided in real time.

[0353] Step 11:

[0354] Users enjoy the viewing experience through the video, audio, and haptic feedback provided by the device. Emotion-based haptic adjustments allow for a more immersive experience.

[0355] (Example 2)

[0356] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0357] Traditional audiovisual content has relied on visual and auditory experiences, but it lacks mechanisms for users to become more emotionally immersed in the content and engage in interactive experiences. In particular, there is a need to enhance the sensory experience by providing tactile feedback in addition to visual and auditory information. Furthermore, a challenge lies in achieving more personalized responses by dynamically adjusting tactile signals according to the user's emotional state.

[0358] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0359] In this invention, the server includes means for using a generation method to analyze audiovisual information, identify specific scenes, and generate acoustic information; means for creating haptic signals from the generated acoustic information; and analysis device means for analyzing the user's emotional state and dynamically adjusting the haptic signals based on that state. This enables the user to obtain an immersive experience that integrates visual, auditory, and tactile senses with the content they are viewing, and in addition, enables personalized haptic feedback that responds to the user's real-time emotional state.

[0360] "Audiovisual information" is a general term for information that includes visual data related to vision and audio data related to hearing.

[0361] A "generation method" is a technical means for identifying specific scenes from audiovisual information and creating audio information tailored to a particular purpose.

[0362] "Acoustic information" refers to sound information, including music and sound effects appropriate to a scene, generated through the analysis of audiovisual information.

[0363] "Haptic signals" are signals that represent physical stimuli such as vibration and pressure, designed to provide haptic feedback to the user.

[0364] An "analysis device" refers to a device or software that has the function of collecting and analyzing real-time data from users and evaluating their emotional state.

[0365] "Combined information" refers to a dataset that integrates visual, auditory, and tactile information, which is transmitted to the user's terminal device to provide a holistic experience.

[0366] A "receiving device" refers to a device that receives composite information transmitted from a server, and is used to provide users with audiovisual experiences and haptic feedback.

[0367] An "output device" refers to a device that transmits physical haptic feedback to the user based on the received haptic signals.

[0368] "Perceptual experience" refers to an interactive experience provided to the user through an integrated sensory experience encompassing sight, hearing, and touch.

[0369] This invention relates to a system for enhancing the user experience based on audiovisual information. This system consists of a server, a terminal, and a user, and each component works in conjunction to provide an unprecedented multi-sensory experience.

[0370] The server receives audiovisual information and analyzes that data. In particular, it uses generation techniques to identify specific scenes from the audiovisual information and generate appropriate audio information. The server can use image analysis software and audio analysis tools for the processing required for analysis. For example, it can use OpenCV, an open-source image processing library, and libraries for extracting audio features.

[0371] The server then creates haptic signals from the generated acoustic information. AI technology is used to design haptic feedback that is synchronized with the acoustic information. Since the haptic signals are physically fed back to the user as vibrations and pressure, their design is crucial.

[0372] Furthermore, the server analyzes the user's emotional state. To do this, it uses an emotion analysis device to analyze the user's real-time data (e.g., camera footage and audio data). This allows the server to dynamically adjust haptic signals in response to changes in the user's emotions. This dynamic adjustment makes it possible to provide the user with a personalized haptic experience.

[0373] The server transmits visual, auditory, and dynamically adjusted haptic signals as a composite information to the terminal. This composite information is transmitted quickly and accurately using real-time communication technology. Low-latency methods such as WebSocket are suitable as the communication protocol.

[0374] The device decomposes the received composite information and reproduces the visual and auditory information on the user's screen or speaker. Meanwhile, haptic signals are used to drive a vibration motor, providing haptic feedback. The vibration output device included in the device delivers accurate haptic feedback to the user based on a pre-designed vibration pattern.

[0375] Users can enjoy an integrated visual, auditory, and tactile experience delivered through the device. This system provides a deeper sense of immersion and a more personalized interactive experience by offering adaptive feedback that responds to the user's emotions.

[0376] As a concrete example of a prompt, one could input the following into the generating AI model: "Explain how to dynamically adjust and display haptic feedback according to the user's emotions." By designing the operation according to this example, it becomes possible to realize the new user experience that the system aims for.

[0377] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0378] Step 1:

[0379] The server receives audiovisual data from video content creators. This data includes video and audio data. The received audiovisual data is converted to a different format by the server and prepared for analysis.

[0380] Step 2:

[0381] The server analyzes audiovisual data to identify specific scenes. This analysis uses image analysis software to analyze video data frame by frame and select specific scenes. It also uses audio analysis tools to extract features from audio data. The input to this process is video data and audio data, and the output is identified scene information and audio feature data.

[0382] Step 3:

[0383] The server generates acoustic information using a generation method based on the identified scene. A generation AI model is utilized to create music and sound effects that are optimal for the selected scene. The input is identified scene information and audio feature data, and the output is the generated acoustic information.

[0384] Step 4:

[0385] The server creates haptic signals based on the generated acoustic information. The haptic signals are data representations of vibration patterns designed to match the acoustic information. In this step, the generated acoustic information is used as input, and haptic signal data is obtained as output.

[0386] Step 5:

[0387] The server collects real-time emotional data from users and performs emotional analysis. This includes analyzing camera footage and audio input to evaluate the user's emotional state. The input is real-time data from the user, and the output is emotional state information obtained through analysis.

[0388] Step 6:

[0389] The server dynamically adjusts haptic signals based on the user's emotional state information. The vibration intensity and pattern of the haptic signals are modified according to the analysis results. The input consists of emotional state information and haptic signal data, while the output is the adjusted haptic signal.

[0390] Step 7:

[0391] The server packages the tuned haptic signals and audiovisual information as composite data and transmits it to the terminal. A low-latency communication protocol is used to transmit the composite data. The input is the tuned haptic signals and audiovisual information, and the output is the data to be transmitted to the terminal.

[0392] Step 8:

[0393] The terminal decomposes the received composite information and reproduces the visual and auditory information for the user. Simultaneously, it transmits haptic signals to a vibration device, providing the user with physical feedback. The input is data transmitted from the server, and the output is video, audio, and haptic feedback.

[0394] (Application Example 2)

[0395] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0396] Modern content delivery services demand immersive experiences that integrate not only audiovisual but also tactile senses. However, current technology makes it difficult to provide dynamic haptic feedback in real time that responds to the user's emotions. Therefore, there are technical challenges to further enhance the viewer's sensory experience.

[0397] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0398] In this invention, the server includes means for using generation technology to analyze audiovisual signals and generate representative acoustic information, means for creating tactile information based on the generated acoustic information, and data processing means for dynamically adjusting the tactile information based on the user's emotions and transmitting the composite information to a terminal device. This makes it possible to provide the user with real-time tactile feedback synchronized with audiovisual data and optimize the perceptual experience.

[0399] "Audiovisual signals" are data that includes image and sound information, and are fundamental elements for users to experience content through sight and hearing.

[0400] "Acoustic information" is data that represents the characteristics of sound, generated by analyzing audiovisual signals, and extracts representative sounds and important sound changes.

[0401] "Haptic information" refers to signals generated to provide feedback that directly appeals to the user's senses, reproducing physical sensations such as vibrations.

[0402] "User emotions" refer to information that indicates the viewer's current feelings and mental state, and are analyzed in real time from facial expressions and biosignals.

[0403] "Composite information" refers to a data package that integrates audiovisual signals, acoustic information, and haptic information, and is transmitted to terminal devices to provide users with an immersive experience.

[0404] "Terminal device" refers to a device used by a user to experience content, and includes smartphones, head-mounted displays, and other similar devices.

[0405] "Data processing" refers to a series of computational processes that analyze, integrate, and transmit multiple pieces of information, and is a central function of a system for optimizing the user experience.

[0406] The system of the present invention aims to provide users with an immersive experience that integrates sight, hearing, and touch. This system mainly consists of two main components: a server and a terminal device. The server receives audiovisual signals, analyzes important scenes using generation technology, and generates representative acoustic information. Subsequently, it creates tactile information based on this acoustic information and dynamically adjusts the tactile information using an emotion engine that analyzes the user's emotions in real time. This adjusted tactile information is packaged together with the audiovisual signals as composite information and transmitted to the terminal device.

[0407] The terminal device receives this composite information, plays back the audio information, and simultaneously materializes the haptic information using a physical device. This allows users to enhance their content experience not only through sight and hearing, but also through touch. Specific hardware includes smartphones and head-mounted displays, while software such as EmotionEngine and HapticFeedback is used.

[0408] As a concrete example, when a user is watching an emotionally moving film, the server analyzes the viewer's facial expressions, captures subtle emotional changes, and adjusts the haptic information accordingly. Soft vibrations are provided during emotional scenes, while stronger vibrations are generated during action scenes, resulting in an interactive experience.

[0409] Another example of providing prompts to a generative AI model is, "Design a system that analyzes user emotions in real time and provides dynamic haptic feedback according to the video scene." Based on this prompt, the AI ​​model generates an appropriate response, enabling the design of enhanced systems.

[0410] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0411] Step 1:

[0412] The server receives audiovisual signals from content creators. These signals include the video and audio that the user will view. The server uses generation technology to analyze these audiovisual signals and identify important scenes. This analysis allows the server to extract video features and acoustic information and generate representative acoustic information corresponding to the important scenes.

[0413] Step 2:

[0414] The server generates haptic information based on the generated acoustic information. This haptic information is designed to provide the user with a sensory experience in the form of vibration and pressure. This process analyzes the intensity and changes in the acoustic information and designs corresponding vibration patterns. As a result, haptic information that the user actually feels is constructed.

[0415] Step 3:

[0416] The server uses an emotion engine to analyze the user's emotions in real time. It receives the user's facial expression data and biosignals as input and evaluates the current emotional state based on that data. The results of the emotion analysis are used to dynamically adjust haptic information. This makes it possible to design optimal haptic feedback that matches the user's emotions.

[0417] Step 4:

[0418] The server integrates haptic information with audiovisual signals and packages it as composite information. This process seamlessly integrates haptic, acoustic, and visual information to generate a dataset that provides the user with a more complete sensory experience. The generated composite information is then ready to be transmitted to the terminal device.

[0419] Step 5:

[0420] The terminal receives composite information transmitted from the server. It decomposes the composite information, plays the acoustic information through an audio output device, and displays the visual information on the display. It also analyzes tactile information and converts it into signals to control a vibration motor, providing the user with physical tactile feedback.

[0421] Step 6:

[0422] Users enjoy a multi-sensory experience through their devices. By experiencing content that integrates sight, hearing, and touch, they can achieve an unprecedented level of immersion. Through this process, users can feel as if they are actually inside the video.

[0423] This series of processes allows for the dynamic adjustment of haptic information based on the user's emotions, while utilizing audiovisual signals, to provide a personalized entertainment experience.

[0424] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0425] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0426] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0427] [Third Embodiment]

[0428] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0429] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0430] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0431] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0432] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0433] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0434] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0435] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0436] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0437] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0438] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0439] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0440] The server receives audiovisual data such as movies, dramas, and documentaries. This data contains information about the images and sounds that viewers experience. The server uses generation technology to analyze this data and generate representative sounds corresponding to specific scenes. These representative sounds refer to sounds that strongly influence both visual and auditory perception, such as explosions in action scenes in movies or melodies in emotional scenes.

[0441] Based on the generated representative sound, the server creates a haptic signal. This haptic signal is intended to provide physical feedback, such as vibration, to the viewer and is transmitted to the user's terminal device. The server then constructs composite data, including the haptic signal, and transmits it to the terminal device through the distribution infrastructure.

[0442] The terminal receives composite data transmitted from the server, decomposes it, and obtains individual signals (video, audio, and haptic signals). The terminal plays back the video and audio signals, providing the user with a visual and auditory experience. The terminal also processes the haptic signals and generates vibrations via an onboard physical device. This feedback is provided in real time, synchronized with the movie scenes.

[0443] Users select video content using their terminal device and receive various media information visually and aurally. Haptic feedback allows users to experience a sense of reality and immersion that cannot be obtained through video and audio alone. For example, when watching an action movie, the device vibrates during explosion scenes, providing a more immersive experience.

[0444] In this way, the present invention aims to realize a new media experience that integrates the three senses of sight, hearing, and touch, and to provide users with unprecedented immersive entertainment.

[0445] The following describes the processing flow.

[0446] Step 1:

[0447] The server receives audiovisual data from video streaming service providers. This data includes video and audio information, which is then analyzed and processed.

[0448] Step 2:

[0449] The server receives audiovisual data and analyzes it using generation technology. Here, it identifies important scenes and generates representative sounds corresponding to them.

[0450] Step 3:

[0451] The server creates haptic signals based on the generated representative sound. This includes setting specific vibration patterns. For example, a strong, rapid vibration might be set for an explosion sound.

[0452] Step 4:

[0453] The server integrates video, audio, and haptic signals to create a composite data package. This data includes visual, auditory, and tactile information.

[0454] Step 5:

[0455] The server sends the generated composite data package to the terminal device. This transmission is performed using the network infrastructure and is delivered in real-time or on-demand.

[0456] Step 6:

[0457] The terminal receives the composite data sent from the server. The data arrives via the communication module and is ready for use.

[0458] Step 7:

[0459] The terminal analyzes the received composite data and separates it into video, audio, and haptic signals. These signals are then distributed to their respective playback systems.

[0460] Step 8:

[0461] The terminal sends video and audio signals to corresponding output devices (displays and speakers) and displays and plays them in a format that the user can view.

[0462] Step 9:

[0463] The device sends haptic signals to a vibration motor, generating physical vibrations in real time. This allows the user to receive haptic feedback.

[0464] Step 10:

[0465] Users enjoy audiovisual information provided through their devices and gain immersion through haptic feedback. Haptic feedback has the effect of amplifying the impact and emotion of the scenes.

[0466] (Example 1)

[0467] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0468] Current audiovisual experiences are limited to sight and hearing, which presents a challenge in that users cannot deeply feel a sense of presence or immersion in the content. This challenge can prevent viewers from fully enjoying the emotional impact and entertainment value of the content.

[0469] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0470] In this invention, the server includes means for using generation technology to analyze audiovisual information and generate characteristic sounds, means for creating vibration signals based on the generated sounds, and information processing means for transmitting composite information including vibration signals to an information processing device. This enables the user to have a new media experience that integrates the three senses of sight, hearing, and touch.

[0471] "Audiovisual information" is a general term for video and audio data acquired through sight and hearing.

[0472] "Distinctive sound" refers to an audio signal generated to emphasize a specific scene within audiovisual information.

[0473] "Generative technology" refers to the technology of analyzing data and creating new acoustic signals.

[0474] A "vibration signal" is a signal used to generate physical vibrations.

[0475] "Complex information" refers to a dataset that integrates video signals, audio signals, vibration signals, and other data.

[0476] An "information processing device" is a device that decomposes received complex information and distributes each signal to the appropriate output device.

[0477] A "physical device" is a device that receives vibration signals and generates actual vibrations.

[0478] "Sensory experience" refers to a new experience that a user gains through sight, hearing, and touch.

[0479] In implementing this invention, the server plays the role of receiving audiovisual information. Specifically, it acquires visual and auditory data from content such as movies, dramas, and documentaries. The server analyzes this data using a generative AI model and generates characteristic sounds associated with a specific scene. An example of a prompt sentence to be input to the AI ​​model is, "Generate representative sounds from an action scene in a movie."

[0480] Based on the generated sound, the server creates a vibration signal. For example, if there is an explosion scene in a movie, a strong vibration signal will be generated. This vibration signal is combined with the video and audio signals as composite information and transmitted from the server to the terminal, which is an information processing device.

[0481] The device decomposes the received composite information and outputs video signals to the display and audio signals to the speaker, providing the user with an audiovisual experience. Furthermore, vibration signals are reproduced as actual vibrations by a physical device built into the device. In this way, the user can obtain a sensory experience that utilizes not only sight and hearing, but also touch. This haptic feedback enhances the sense of presence and immersion in the content, providing a richer entertainment experience.

[0482] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0483] Step 1:

[0484] The server receives audiovisual information. This data includes video and audio data from movies, dramas, and documentaries. This data is used as input for analysis by a generative AI model.

[0485] Step 2:

[0486] The server analyzes the input audiovisual information using a generative AI model. The AI ​​model receives the prompt "Generate representative sounds from an action scene in a movie," and based on this instruction, identifies a specific scene from the audiovisual data and generates characteristic sounds. The output is characteristic sound data for that scene.

[0487] Step 3:

[0488] The server receives the generated acoustic data as input and creates a vibration signal based on it. For example, if the generated sound is an explosion, a strong vibration signal is created. This vibration signal is output as partial data for later composite information construction.

[0489] Step 4:

[0490] The server integrates video, audio, and vibration signals to form composite information. This composite information becomes output data to be transmitted to the terminal, which is an information processing device.

[0491] Step 5:

[0492] The terminal receives composite information transmitted from the server. This information includes video signals, audio signals, and vibration signals, which are then separated and used as input data to be appropriately distributed to each output device.

[0493] Step 6:

[0494] The device outputs the decomposed video signal to the display and the audio signal to the speaker. This provides the user with an audiovisual experience. The output here is the actual content that the user experiences visually and aurally.

[0495] Step 7:

[0496] The device receives vibration signals and reproduces actual vibrations using its built-in physical components. The device then delivers physical vibrations as output, allowing the user to experience realistic feedback through touch. This process enables the user to obtain an immersive and realistic media experience.

[0497] (Application Example 1)

[0498] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0499] Modern audiovisual content, which relies solely on video and audio for the experience, limits interaction with the real world and a sense of presence. There is a need for methods to provide users with more immersive and sensory-rich experiences.

[0500] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0501] In this invention, the server includes means for analyzing audiovisual information and using a generation method to generate characteristic sounds, means for generating tactile signals from the generated characteristic sounds, and information processing means for transmitting composite information including these tactile signals to an information terminal. This makes it possible to improve the user's sensory experience by providing tactile feedback.

[0502] "Audiovisual information" refers to digital data related to sight and hearing, including images and sounds.

[0503] A "characteristic sound" is an acoustic signal that represents a specific scene within audiovisual information and is generated to enhance the user experience.

[0504] A "generation method" is a computational procedure or algorithm used to achieve a specific purpose, and is a technique used to generate characteristic sounds from audiovisual information.

[0505] "Haptic signals" are signals that elicit physical feedback, such as vibration or pressure, and provide users with tactile interaction.

[0506] "Composite information" refers to a collection of information including images, sounds, and haptic signals, which is transmitted to an information terminal to provide users with a multi-sensory experience.

[0507] An "information terminal" is an electronic device equipped with the function of displaying and playing content, which processes the received composite information to provide an experience to the user.

[0508] The system implementing this invention mainly consists of a server and an information terminal. The server analyzes audiovisual information using a generative AI model and generates characteristic sounds that represent a specific scene. Based on these characteristic sounds, the server generates tactile signals and transmits them to the information terminal as composite information.

[0509] The information terminal decomposes the received composite information and plays back video and audio to provide the user with an audiovisual experience. Furthermore, it realizes haptic signals through a physical device, providing the user with real-time haptic feedback. This feedback allows the user to experience a sense of presence that cannot be obtained from video and audio alone.

[0510] Implementing this system requires information terminals such as smartphones or tablets equipped with vibration motors. The server software uses a generative AI model to analyze audiovisual information and generate characteristic sounds. In particular, APIs for generating tactile signals, such as OpenHaptics, are utilized.

[0511] For example, during an action scene in a movie, the smartphone might vibrate strongly, allowing the user to feel the tension of that scene. This kind of feedback provides a new entertainment experience that integrates not only sight and hearing but also touch.

[0512] An example of a prompt message sent to a generative AI model is: "The next scene is an emotional climax. Please generate subtle haptic feedback that will evoke emotions in the user."

[0513] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0514] Step 1:

[0515] The server receives audiovisual information. It takes video and audio data from movies and dramas as input and analyzes this data using a generative AI model. As a result of the analysis, it outputs characteristic sounds that represent a specific scene.

[0516] Step 2:

[0517] The server generates haptic signals based on characteristic sounds. It takes characteristic sounds as input and outputs the generated haptic signals. Here, vibration patterns are designed using APIs such as OpenHaptics to generate haptic signals suitable for feedback.

[0518] Step 3:

[0519] The server combines the generated haptic signals with video and audio data to form composite information. This composite information is the data output and is prepared for transmission to information terminals.

[0520] Step 4:

[0521] The information terminal decomposes the complex information received from the server. It separates the complex information received as input into video signals, audio signals, and haptic signals, and processes each of these individually.

[0522] Step 5:

[0523] The information terminal plays video and audio signals, providing the user with an audiovisual experience. In this step, the media player on the information terminal operates, providing the user with audiovisual content from video and audio.

[0524] Step 6:

[0525] The information terminal executes haptic signals via a physical device. Here, haptic signals are input to a physical device such as a vibration motor, and a specified vibration feedback is provided to the user as output.

[0526] Step 7:

[0527] Users enjoy a tripartite sensory experience—audiovisual and tactile—provided by the information terminal. In this way, they experience immersive entertainment with haptic feedback optimized for each scene in movies and dramas.

[0528] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0529] The server receives audiovisual data from video content creators, which includes video and audio information. First, the server analyzes this audiovisual data using generation technology. It identifies important scenes and events and generates corresponding representative sounds. By creating haptic signals based on the generated representative sounds, it prepares to provide users with haptic feedback that is linked to their audiovisual experience.

[0530] The server further utilizes an emotion engine to analyze the user's emotions in real time. The emotion engine uses data such as the user's facial expressions, heart rate, and speech to understand their current emotional state. Based on this emotional information, the server dynamically adjusts haptic signals to create haptic feedback that is optimal for the user's current emotions.

[0531] The generated haptic signals are packaged together with visual and auditory signals as composite data and transmitted to a terminal device. The terminal receives this composite data and plays the video and audio signals, while also processing the haptic signals and implementing them with a vibration motor. This allows the user to experience a deeper sense of immersion in the content they are viewing.

[0532] For example, when a user is watching an emotionally moving film, the emotion engine detects the user's tears or smiles and adjusts the haptic signals to a milder, more pleasant vibration. Similarly, during action scenes, it senses the user's excitement level and provides strong, vivid vibration feedback. This real-time emotional feedback allows users to enjoy a more interactive and personalized entertainment experience.

[0533] This technology expands the possibilities of unprecedented viewing experiences by integrating visual, auditory, and tactile senses, as well as providing adaptive feedback tailored to the user's emotional state. Seamless communication between the server and the device, along with the recognition and application of emotions, allows users to obtain individually optimized content experiences.

[0534] The following describes the processing flow.

[0535] Step 1:

[0536] The server receives audiovisual data, including video and audio, from video streaming providers. The received data is prepared for analysis.

[0537] Step 2:

[0538] The server uses generation technology to analyze audiovisual data and generate representative sounds corresponding to specific scenes. This analysis identifies key events in the video and generates acoustic information based on them.

[0539] Step 3:

[0540] The server creates a haptic signal based on the generated representative sound. This haptic signal includes an appropriate vibration pattern corresponding to the scene of the content being viewed.

[0541] Step 4:

[0542] The server uses an emotion engine to analyze the user's emotions in real time. Based on sensor data received from the user (e.g., camera footage and heart rate), it analyzes the user's emotional state and creates an emotion profile.

[0543] Step 5:

[0544] The server adjusts haptic signals based on the user's emotional profile. This dynamically changes the intensity and pattern of vibrations according to the user's emotional state to provide optimal feedback.

[0545] Step 6:

[0546] The server creates a composite data package containing video, audio, and tuned haptic signals. This package is then prepared in the appropriate format for transmission to the terminal device.

[0547] Step 7:

[0548] The server transmits the composite data package to the terminal device. The transmission takes place over the network, ensuring that the data is transmitted correctly and without delay.

[0549] Step 8:

[0550] The terminal receives the composite data transmitted from the server. After receiving, the data is analyzed for processing and separated into its respective signal formats.

[0551] Step 9:

[0552] The terminal transmits separated video and audio signals to the playback device, providing the user with visual and auditory feedback.

[0553] Step 10:

[0554] The device transmits haptic signals to a vibration motor, generating vibrations that allow the user to experience haptic feedback. This feedback is provided in real time.

[0555] Step 11:

[0556] Users enjoy the viewing experience through the video, audio, and haptic feedback provided by the device. Emotion-based haptic adjustments allow for a more immersive experience.

[0557] (Example 2)

[0558] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0559] Traditional audiovisual content has relied on visual and auditory experiences, but it lacks mechanisms for users to become more emotionally immersed in the content and engage in interactive experiences. In particular, there is a need to enhance the sensory experience by providing tactile feedback in addition to visual and auditory information. Furthermore, a challenge lies in achieving more personalized responses by dynamically adjusting tactile signals according to the user's emotional state.

[0560] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0561] In this invention, the server includes means for using a generation method to analyze audiovisual information, identify specific scenes, and generate acoustic information; means for creating haptic signals from the generated acoustic information; and analysis device means for analyzing the user's emotional state and dynamically adjusting the haptic signals based on that state. This enables the user to obtain an immersive experience that integrates visual, auditory, and tactile senses with the content they are viewing, and in addition, enables personalized haptic feedback that responds to the user's real-time emotional state.

[0562] "Audiovisual information" is a general term for information that includes visual data related to vision and audio data related to hearing.

[0563] A "generation method" is a technical means for identifying specific scenes from audiovisual information and creating audio information tailored to a particular purpose.

[0564] "Acoustic information" refers to sound information, including music and sound effects appropriate to a scene, generated through the analysis of audiovisual information.

[0565] "Haptic signals" are signals that represent physical stimuli such as vibration and pressure, designed to provide haptic feedback to the user.

[0566] An "analysis device" refers to a device or software that has the function of collecting and analyzing real-time data from users and evaluating their emotional state.

[0567] "Combined information" refers to a dataset that integrates visual, auditory, and tactile information, which is transmitted to the user's terminal device to provide a holistic experience.

[0568] A "receiving device" refers to a device that receives composite information transmitted from a server, and is used to provide users with audiovisual experiences and haptic feedback.

[0569] An "output device" refers to a device that transmits physical haptic feedback to the user based on the received haptic signals.

[0570] "Perceptual experience" refers to an interactive experience provided to the user through an integrated sensory experience encompassing sight, hearing, and touch.

[0571] This invention relates to a system for enhancing the user experience based on audiovisual information. This system consists of a server, a terminal, and a user, and each component works in conjunction to provide an unprecedented multi-sensory experience.

[0572] The server receives audiovisual information and analyzes that data. In particular, it uses generation techniques to identify specific scenes from the audiovisual information and generate appropriate audio information. The server can use image analysis software and audio analysis tools for the processing required for analysis. For example, it can use OpenCV, an open-source image processing library, and libraries for extracting audio features.

[0573] The server then creates haptic signals from the generated acoustic information. AI technology is used to design haptic feedback that is synchronized with the acoustic information. Since the haptic signals are physically fed back to the user as vibrations and pressure, their design is crucial.

[0574] Furthermore, the server analyzes the user's emotional state. To do this, it uses an emotion analysis device to analyze the user's real-time data (e.g., camera footage and audio data). This allows the server to dynamically adjust haptic signals in response to changes in the user's emotions. This dynamic adjustment makes it possible to provide the user with a personalized haptic experience.

[0575] The server transmits visual, auditory, and dynamically adjusted haptic signals as a composite information to the terminal. This composite information is transmitted quickly and accurately using real-time communication technology. Low-latency methods such as WebSocket are suitable as the communication protocol.

[0576] The device decomposes the received composite information and reproduces the visual and auditory information on the user's screen or speaker. Meanwhile, haptic signals are used to drive a vibration motor, providing haptic feedback. The vibration output device included in the device delivers accurate haptic feedback to the user based on a pre-designed vibration pattern.

[0577] Users can enjoy an integrated visual, auditory, and tactile experience delivered through the device. This system provides a deeper sense of immersion and a more personalized interactive experience by offering adaptive feedback that responds to the user's emotions.

[0578] As a concrete example of a prompt, one could input the following into the generating AI model: "Explain how to dynamically adjust and display haptic feedback according to the user's emotions." By designing the operation according to this example, it becomes possible to realize the new user experience that the system aims for.

[0579] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0580] Step 1:

[0581] The server receives audiovisual data from video content creators. This data includes video and audio data. The received audiovisual data is converted to a different format by the server and prepared for analysis.

[0582] Step 2:

[0583] The server analyzes audiovisual data to identify specific scenes. This analysis uses image analysis software to analyze video data frame by frame and select specific scenes. It also uses audio analysis tools to extract features from audio data. The input to this process is video data and audio data, and the output is identified scene information and audio feature data.

[0584] Step 3:

[0585] The server generates acoustic information using a generation method based on the identified scene. A generation AI model is utilized to create music and sound effects that are optimal for the selected scene. The input is identified scene information and audio feature data, and the output is the generated acoustic information.

[0586] Step 4:

[0587] The server creates haptic signals based on the generated acoustic information. The haptic signals are data representations of vibration patterns designed to match the acoustic information. In this step, the generated acoustic information is used as input, and haptic signal data is obtained as output.

[0588] Step 5:

[0589] The server collects real-time emotional data from users and performs emotional analysis. This includes analyzing camera footage and audio input to evaluate the user's emotional state. The input is real-time data from the user, and the output is emotional state information obtained through analysis.

[0590] Step 6:

[0591] The server dynamically adjusts haptic signals based on the user's emotional state information. The vibration intensity and pattern of the haptic signals are modified according to the analysis results. The input consists of emotional state information and haptic signal data, while the output is the adjusted haptic signal.

[0592] Step 7:

[0593] The server packages the tuned haptic signals and audiovisual information as composite data and transmits it to the terminal. A low-latency communication protocol is used to transmit the composite data. The input is the tuned haptic signals and audiovisual information, and the output is the data to be transmitted to the terminal.

[0594] Step 8:

[0595] The terminal decomposes the received composite information and reproduces the visual and auditory information for the user. Simultaneously, it transmits haptic signals to a vibration device, providing the user with physical feedback. The input is data transmitted from the server, and the output is video, audio, and haptic feedback.

[0596] (Application Example 2)

[0597] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0598] Modern content delivery services demand immersive experiences that integrate not only audiovisual but also tactile senses. However, current technology makes it difficult to provide dynamic haptic feedback in real time that responds to the user's emotions. Therefore, there are technical challenges to further enhance the viewer's sensory experience.

[0599] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0600] In this invention, the server includes means for using generation technology to analyze audiovisual signals and generate representative acoustic information, means for creating tactile information based on the generated acoustic information, and data processing means for dynamically adjusting the tactile information based on the user's emotions and transmitting the composite information to a terminal device. This makes it possible to provide the user with real-time tactile feedback synchronized with audiovisual data and optimize the perceptual experience.

[0601] "Audiovisual signals" are data that includes image and sound information, and are fundamental elements for users to experience content through sight and hearing.

[0602] "Acoustic information" is data that represents the characteristics of sound, generated by analyzing audiovisual signals, and extracts representative sounds and important sound changes.

[0603] "Haptic information" refers to signals generated to provide feedback that directly appeals to the user's senses, reproducing physical sensations such as vibrations.

[0604] "User emotions" refer to information that indicates the viewer's current feelings and mental state, and are analyzed in real time from facial expressions and biosignals.

[0605] "Composite information" refers to a data package that integrates audiovisual signals, acoustic information, and haptic information, and is transmitted to terminal devices to provide users with an immersive experience.

[0606] "Terminal device" refers to a device used by a user to experience content, and includes smartphones, head-mounted displays, and other similar devices.

[0607] "Data processing" refers to a series of computational processes that analyze, integrate, and transmit multiple pieces of information, and is a central function of a system for optimizing the user experience.

[0608] The system of the present invention aims to provide users with an immersive experience that integrates sight, hearing, and touch. This system mainly consists of two main components: a server and a terminal device. The server receives audiovisual signals, analyzes important scenes using generation technology, and generates representative acoustic information. Subsequently, it creates tactile information based on this acoustic information and dynamically adjusts the tactile information using an emotion engine that analyzes the user's emotions in real time. This adjusted tactile information is packaged together with the audiovisual signals as composite information and transmitted to the terminal device.

[0609] The terminal device receives this composite information, plays back the audio information, and simultaneously materializes the haptic information using a physical device. This allows users to enhance their content experience not only through sight and hearing, but also through touch. Specific hardware includes smartphones and head-mounted displays, while software such as EmotionEngine and HapticFeedback is used.

[0610] As a concrete example, when a user is watching an emotionally moving film, the server analyzes the viewer's facial expressions, captures subtle emotional changes, and adjusts the haptic information accordingly. Soft vibrations are provided during emotional scenes, while stronger vibrations are generated during action scenes, resulting in an interactive experience.

[0611] Another example of providing prompts to a generative AI model is, "Design a system that analyzes user emotions in real time and provides dynamic haptic feedback according to the video scene." Based on this prompt, the AI ​​model generates an appropriate response, enabling the design of enhanced systems.

[0612] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0613] Step 1:

[0614] The server receives audiovisual signals from content creators. These signals include the video and audio that the user will view. The server uses generation technology to analyze these audiovisual signals and identify important scenes. This analysis allows the server to extract video features and acoustic information and generate representative acoustic information corresponding to the important scenes.

[0615] Step 2:

[0616] The server generates haptic information based on the generated acoustic information. This haptic information is designed to provide the user with a sensory experience in the form of vibration and pressure. This process analyzes the intensity and changes in the acoustic information and designs corresponding vibration patterns. As a result, haptic information that the user actually feels is constructed.

[0617] Step 3:

[0618] The server uses an emotion engine to analyze the user's emotions in real time. It receives the user's facial expression data and biosignals as input and evaluates the current emotional state based on that data. The results of the emotion analysis are used to dynamically adjust haptic information. This makes it possible to design optimal haptic feedback that matches the user's emotions.

[0619] Step 4:

[0620] The server integrates haptic information with audiovisual signals and packages it as composite information. This process seamlessly integrates haptic, acoustic, and visual information to generate a dataset that provides the user with a more complete sensory experience. The generated composite information is then ready to be transmitted to the terminal device.

[0621] Step 5:

[0622] The terminal receives composite information transmitted from the server. It decomposes the composite information, plays the acoustic information through an audio output device, and displays the visual information on the display. It also analyzes tactile information and converts it into signals to control a vibration motor, providing the user with physical tactile feedback.

[0623] Step 6:

[0624] Users enjoy a multi-sensory experience through their devices. By experiencing content that integrates sight, hearing, and touch, they can achieve an unprecedented level of immersion. Through this process, users can feel as if they are actually inside the video.

[0625] This series of processes allows for the dynamic adjustment of haptic information based on the user's emotions, while utilizing audiovisual signals, to provide a personalized entertainment experience.

[0626] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0627] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0628] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0629] [Fourth Embodiment]

[0630] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0631] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0632] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0633] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0634] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0635] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0636] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0637] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0638] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0639] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0640] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0641] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0642] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0643] The server receives audiovisual data such as movies, dramas, and documentaries. This data contains information about the images and sounds that viewers experience. The server uses generation technology to analyze this data and generate representative sounds corresponding to specific scenes. These representative sounds refer to sounds that strongly influence both visual and auditory perception, such as explosions in action scenes in movies or melodies in emotional scenes.

[0644] Based on the generated representative sound, the server creates a haptic signal. This haptic signal is intended to provide physical feedback, such as vibration, to the viewer and is transmitted to the user's terminal device. The server then constructs composite data, including the haptic signal, and transmits it to the terminal device through the distribution infrastructure.

[0645] The terminal receives composite data transmitted from the server, decomposes it, and obtains individual signals (video, audio, and haptic signals). The terminal plays back the video and audio signals, providing the user with a visual and auditory experience. The terminal also processes the haptic signals and generates vibrations via an onboard physical device. This feedback is provided in real time, synchronized with the movie scenes.

[0646] Users select video content using their terminal device and receive various media information visually and aurally. Haptic feedback allows users to experience a sense of reality and immersion that cannot be obtained through video and audio alone. For example, when watching an action movie, the device vibrates during explosion scenes, providing a more immersive experience.

[0647] In this way, the present invention aims to realize a new media experience that integrates the three senses of sight, hearing, and touch, and to provide users with unprecedented immersive entertainment.

[0648] The following describes the processing flow.

[0649] Step 1:

[0650] The server receives audiovisual data from video streaming service providers. This data includes video and audio information, which is then analyzed and processed.

[0651] Step 2:

[0652] The server receives audiovisual data and analyzes it using generation technology. Here, it identifies important scenes and generates representative sounds corresponding to them.

[0653] Step 3:

[0654] The server creates haptic signals based on the generated representative sound. This includes setting specific vibration patterns. For example, a strong, rapid vibration might be set for an explosion sound.

[0655] Step 4:

[0656] The server integrates video, audio, and haptic signals to create a composite data package. This data includes visual, auditory, and tactile information.

[0657] Step 5:

[0658] The server sends the generated composite data package to the terminal device. This transmission is performed using the network infrastructure and is delivered in real-time or on-demand.

[0659] Step 6:

[0660] The terminal receives the composite data sent from the server. The data arrives via the communication module and is ready for use.

[0661] Step 7:

[0662] The terminal analyzes the received composite data and separates it into video, audio, and haptic signals. These signals are then distributed to their respective playback systems.

[0663] Step 8:

[0664] The terminal sends video and audio signals to corresponding output devices (displays and speakers) and displays and plays them in a format that the user can view.

[0665] Step 9:

[0666] The device sends haptic signals to a vibration motor, generating physical vibrations in real time. This allows the user to receive haptic feedback.

[0667] Step 10:

[0668] Users enjoy audiovisual information provided through their devices and gain immersion through haptic feedback. Haptic feedback has the effect of amplifying the impact and emotion of the scenes.

[0669] (Example 1)

[0670] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0671] Current audiovisual experiences are limited to sight and hearing, which presents a challenge in that users cannot deeply feel a sense of presence or immersion in the content. This challenge can prevent viewers from fully enjoying the emotional impact and entertainment value of the content.

[0672] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0673] In this invention, the server includes means for using generation technology to analyze audiovisual information and generate characteristic sounds, means for creating vibration signals based on the generated sounds, and information processing means for transmitting composite information including vibration signals to an information processing device. This enables the user to have a new media experience that integrates the three senses of sight, hearing, and touch.

[0674] "Audiovisual information" is a general term for video and audio data acquired through sight and hearing.

[0675] "Distinctive sound" refers to an audio signal generated to emphasize a specific scene within audiovisual information.

[0676] "Generative technology" refers to the technology of analyzing data and creating new acoustic signals.

[0677] A "vibration signal" is a signal used to generate physical vibrations.

[0678] "Complex information" refers to a dataset that integrates video signals, audio signals, vibration signals, and other data.

[0679] An "information processing device" is a device that decomposes received complex information and distributes each signal to the appropriate output device.

[0680] A "physical device" is a device that receives vibration signals and generates actual vibrations.

[0681] "Sensory experience" refers to a new experience that a user gains through sight, hearing, and touch.

[0682] In implementing this invention, the server plays the role of receiving audiovisual information. Specifically, it acquires visual and auditory data from content such as movies, dramas, and documentaries. The server analyzes this data using a generative AI model and generates characteristic sounds associated with a specific scene. An example of a prompt sentence to be input to the AI ​​model is, "Generate representative sounds from an action scene in a movie."

[0683] Based on the generated sound, the server creates a vibration signal. For example, if there is an explosion scene in a movie, a strong vibration signal will be generated. This vibration signal is combined with the video and audio signals as composite information and transmitted from the server to the terminal, which is an information processing device.

[0684] The device decomposes the received composite information and outputs video signals to the display and audio signals to the speaker, providing the user with an audiovisual experience. Furthermore, vibration signals are reproduced as actual vibrations by a physical device built into the device. In this way, the user can obtain a sensory experience that utilizes not only sight and hearing, but also touch. This haptic feedback enhances the sense of presence and immersion in the content, providing a richer entertainment experience.

[0685] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0686] Step 1:

[0687] The server receives audiovisual information. This data includes video and audio data from movies, dramas, and documentaries. This data is used as input for analysis by a generative AI model.

[0688] Step 2:

[0689] The server analyzes the input audiovisual information using a generative AI model. The AI ​​model receives the prompt "Generate representative sounds from an action scene in a movie," and based on this instruction, identifies a specific scene from the audiovisual data and generates characteristic sounds. The output is characteristic sound data for that scene.

[0690] Step 3:

[0691] The server receives the generated acoustic data as input and creates a vibration signal based on it. For example, if the generated sound is an explosion, a strong vibration signal is created. This vibration signal is output as partial data for later composite information construction.

[0692] Step 4:

[0693] The server integrates video, audio, and vibration signals to form composite information. This composite information becomes output data to be transmitted to the terminal, which is an information processing device.

[0694] Step 5:

[0695] The terminal receives composite information transmitted from the server. This information includes video signals, audio signals, and vibration signals, which are then separated and used as input data to be appropriately distributed to each output device.

[0696] Step 6:

[0697] The device outputs the decomposed video signal to the display and the audio signal to the speaker. This provides the user with an audiovisual experience. The output here is the actual content that the user experiences visually and aurally.

[0698] Step 7:

[0699] The device receives vibration signals and reproduces actual vibrations using its built-in physical components. The device then delivers physical vibrations as output, allowing the user to experience realistic feedback through touch. This process enables the user to obtain an immersive and realistic media experience.

[0700] (Application Example 1)

[0701] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0702] Modern audiovisual content, which relies solely on video and audio for the experience, limits interaction with the real world and a sense of presence. There is a need for methods to provide users with more immersive and sensory-rich experiences.

[0703] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0704] In this invention, the server includes means for analyzing audiovisual information and using a generation method to generate characteristic sounds, means for generating tactile signals from the generated characteristic sounds, and information processing means for transmitting composite information including these tactile signals to an information terminal. This makes it possible to improve the user's sensory experience by providing tactile feedback.

[0705] "Audiovisual information" refers to digital data related to sight and hearing, including images and sounds.

[0706] A "characteristic sound" is an acoustic signal that represents a specific scene within audiovisual information and is generated to enhance the user experience.

[0707] A "generation method" is a computational procedure or algorithm used to achieve a specific purpose, and is a technique used to generate characteristic sounds from audiovisual information.

[0708] "Haptic signals" are signals that elicit physical feedback, such as vibration or pressure, and provide users with tactile interaction.

[0709] "Composite information" refers to a collection of information including images, sounds, and haptic signals, which is transmitted to an information terminal to provide users with a multi-sensory experience.

[0710] An "information terminal" is an electronic device equipped with the function of displaying and playing content, which processes the received composite information to provide an experience to the user.

[0711] The system implementing this invention mainly consists of a server and an information terminal. The server analyzes audiovisual information using a generative AI model and generates characteristic sounds that represent a specific scene. Based on these characteristic sounds, the server generates tactile signals and transmits them to the information terminal as composite information.

[0712] The information terminal decomposes the received composite information and plays back video and audio to provide the user with an audiovisual experience. Furthermore, it realizes haptic signals through a physical device, providing the user with real-time haptic feedback. This feedback allows the user to experience a sense of presence that cannot be obtained from video and audio alone.

[0713] Implementing this system requires information terminals such as smartphones or tablets equipped with vibration motors. The server software uses a generative AI model to analyze audiovisual information and generate characteristic sounds. In particular, APIs for generating tactile signals, such as OpenHaptics, are utilized.

[0714] For example, during an action scene in a movie, the smartphone might vibrate strongly, allowing the user to feel the tension of that scene. This kind of feedback provides a new entertainment experience that integrates not only sight and hearing but also touch.

[0715] An example of a prompt message sent to a generative AI model is: "The next scene is an emotional climax. Please generate subtle haptic feedback that will evoke emotions in the user."

[0716] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0717] Step 1:

[0718] The server receives audiovisual information. It takes video and audio data from movies and dramas as input and analyzes this data using a generative AI model. As a result of the analysis, it outputs characteristic sounds that represent a specific scene.

[0719] Step 2:

[0720] The server generates haptic signals based on characteristic sounds. It takes characteristic sounds as input and outputs the generated haptic signals. Here, vibration patterns are designed using APIs such as OpenHaptics to generate haptic signals suitable for feedback.

[0721] Step 3:

[0722] The server combines the generated haptic signals with video and audio data to form composite information. This composite information is the data output and is prepared for transmission to information terminals.

[0723] Step 4:

[0724] The information terminal decomposes the complex information received from the server. It separates the complex information received as input into video signals, audio signals, and haptic signals, and processes each of these individually.

[0725] Step 5:

[0726] The information terminal plays video and audio signals, providing the user with an audiovisual experience. In this step, the media player on the information terminal operates, providing the user with audiovisual content from video and audio.

[0727] Step 6:

[0728] The information terminal executes haptic signals via a physical device. Here, haptic signals are input to a physical device such as a vibration motor, and a specified vibration feedback is provided to the user as output.

[0729] Step 7:

[0730] Users enjoy a tripartite sensory experience—audiovisual and tactile—provided by the information terminal. In this way, they experience immersive entertainment with haptic feedback optimized for each scene in movies and dramas.

[0731] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0732] The server receives audiovisual data from video content creators, which includes video and audio information. First, the server analyzes this audiovisual data using generation technology. It identifies important scenes and events and generates corresponding representative sounds. By creating haptic signals based on the generated representative sounds, it prepares to provide users with haptic feedback that is linked to their audiovisual experience.

[0733] The server further utilizes an emotion engine to analyze the user's emotions in real time. The emotion engine uses data such as the user's facial expressions, heart rate, and speech to understand their current emotional state. Based on this emotional information, the server dynamically adjusts haptic signals to create haptic feedback that is optimal for the user's current emotions.

[0734] The generated haptic signals are packaged together with visual and auditory signals as composite data and transmitted to a terminal device. The terminal receives this composite data and plays the video and audio signals, while also processing the haptic signals and implementing them with a vibration motor. This allows the user to experience a deeper sense of immersion in the content they are viewing.

[0735] For example, when a user is watching an emotionally moving film, the emotion engine detects the user's tears or smiles and adjusts the haptic signals to a milder, more pleasant vibration. Similarly, during action scenes, it senses the user's excitement level and provides strong, vivid vibration feedback. This real-time emotional feedback allows users to enjoy a more interactive and personalized entertainment experience.

[0736] This technology expands the possibilities of unprecedented viewing experiences by integrating visual, auditory, and tactile senses, as well as providing adaptive feedback tailored to the user's emotional state. Seamless communication between the server and the device, along with the recognition and application of emotions, allows users to obtain individually optimized content experiences.

[0737] The following describes the processing flow.

[0738] Step 1:

[0739] The server receives audiovisual data, including video and audio, from video streaming providers. The received data is prepared for analysis.

[0740] Step 2:

[0741] The server uses generation technology to analyze audiovisual data and generate representative sounds corresponding to specific scenes. This analysis identifies key events in the video and generates acoustic information based on them.

[0742] Step 3:

[0743] The server creates a haptic signal based on the generated representative sound. This haptic signal includes an appropriate vibration pattern corresponding to the scene of the content being viewed.

[0744] Step 4:

[0745] The server uses an emotion engine to analyze the user's emotions in real time. Based on sensor data received from the user (e.g., camera footage and heart rate), it analyzes the user's emotional state and creates an emotion profile.

[0746] Step 5:

[0747] The server adjusts haptic signals based on the user's emotional profile. This dynamically changes the intensity and pattern of vibrations according to the user's emotional state to provide optimal feedback.

[0748] Step 6:

[0749] The server creates a composite data package containing video, audio, and tuned haptic signals. This package is then prepared in the appropriate format for transmission to the terminal device.

[0750] Step 7:

[0751] The server transmits the composite data package to the terminal device. The transmission takes place over the network, ensuring that the data is transmitted correctly and without delay.

[0752] Step 8:

[0753] The terminal receives the composite data transmitted from the server. After receiving, the data is analyzed for processing and separated into its respective signal formats.

[0754] Step 9:

[0755] The terminal transmits separated video and audio signals to the playback device, providing the user with visual and auditory feedback.

[0756] Step 10:

[0757] The device transmits haptic signals to a vibration motor, generating vibrations that allow the user to experience haptic feedback. This feedback is provided in real time.

[0758] Step 11:

[0759] Users enjoy the viewing experience through the video, audio, and haptic feedback provided by the device. Emotion-based haptic adjustments allow for a more immersive experience.

[0760] (Example 2)

[0761] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0762] Traditional audiovisual content has relied on visual and auditory experiences, but it lacks mechanisms for users to become more emotionally immersed in the content and engage in interactive experiences. In particular, there is a need to enhance the sensory experience by providing tactile feedback in addition to visual and auditory information. Furthermore, a challenge lies in achieving more personalized responses by dynamically adjusting tactile signals according to the user's emotional state.

[0763] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0764] In this invention, the server includes means for using a generation method to analyze audiovisual information, identify specific scenes, and generate acoustic information; means for creating haptic signals from the generated acoustic information; and analysis device means for analyzing the user's emotional state and dynamically adjusting the haptic signals based on that state. This enables the user to obtain an immersive experience that integrates visual, auditory, and tactile senses with the content they are viewing, and in addition, enables personalized haptic feedback that responds to the user's real-time emotional state.

[0765] "Audiovisual information" is a general term for information that includes visual data related to vision and audio data related to hearing.

[0766] A "generation method" is a technical means for identifying specific scenes from audiovisual information and creating audio information tailored to a particular purpose.

[0767] "Acoustic information" refers to sound information, including music and sound effects appropriate to a scene, generated through the analysis of audiovisual information.

[0768] "Haptic signals" are signals that represent physical stimuli such as vibration and pressure, designed to provide haptic feedback to the user.

[0769] An "analysis device" refers to a device or software that has the function of collecting and analyzing real-time data from users and evaluating their emotional state.

[0770] "Combined information" refers to a dataset that integrates visual, auditory, and tactile information, which is transmitted to the user's terminal device to provide a holistic experience.

[0771] A "receiving device" refers to a device that receives composite information transmitted from a server, and is used to provide users with audiovisual experiences and haptic feedback.

[0772] An "output device" refers to a device that transmits physical haptic feedback to the user based on the received haptic signals.

[0773] "Perceptual experience" refers to an interactive experience provided to the user through an integrated sensory experience encompassing sight, hearing, and touch.

[0774] This invention relates to a system for enhancing the user experience based on audiovisual information. This system consists of a server, a terminal, and a user, and each component works in conjunction to provide an unprecedented multi-sensory experience.

[0775] The server receives audiovisual information and analyzes that data. In particular, it uses generation techniques to identify specific scenes from the audiovisual information and generate appropriate audio information. The server can use image analysis software and audio analysis tools for the processing required for analysis. For example, it can use OpenCV, an open-source image processing library, and libraries for extracting audio features.

[0776] The server then creates haptic signals from the generated acoustic information. AI technology is used to design haptic feedback that is synchronized with the acoustic information. Since the haptic signals are physically fed back to the user as vibrations and pressure, their design is crucial.

[0777] Furthermore, the server analyzes the user's emotional state. To do this, it uses an emotion analysis device to analyze the user's real-time data (e.g., camera footage and audio data). This allows the server to dynamically adjust haptic signals in response to changes in the user's emotions. This dynamic adjustment makes it possible to provide the user with a personalized haptic experience.

[0778] The server transmits visual, auditory, and dynamically adjusted haptic signals as a composite information to the terminal. This composite information is transmitted quickly and accurately using real-time communication technology. Low-latency methods such as WebSocket are suitable as the communication protocol.

[0779] The device decomposes the received composite information and reproduces the visual and auditory information on the user's screen or speaker. Meanwhile, haptic signals are used to drive a vibration motor, providing haptic feedback. The vibration output device included in the device delivers accurate haptic feedback to the user based on a pre-designed vibration pattern.

[0780] Users can enjoy an integrated visual, auditory, and tactile experience delivered through the device. This system provides a deeper sense of immersion and a more personalized interactive experience by offering adaptive feedback that responds to the user's emotions.

[0781] As a concrete example of a prompt, one could input the following into the generating AI model: "Explain how to dynamically adjust and display haptic feedback according to the user's emotions." By designing the operation according to this example, it becomes possible to realize the new user experience that the system aims for.

[0782] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0783] Step 1:

[0784] The server receives audiovisual data from video content creators. This data includes video and audio data. The received audiovisual data is converted to a different format by the server and prepared for analysis.

[0785] Step 2:

[0786] The server analyzes audiovisual data to identify specific scenes. This analysis uses image analysis software to analyze video data frame by frame and select specific scenes. It also uses audio analysis tools to extract features from audio data. The input to this process is video data and audio data, and the output is identified scene information and audio feature data.

[0787] Step 3:

[0788] The server generates acoustic information using a generation method based on the identified scene. A generation AI model is utilized to create music and sound effects that are optimal for the selected scene. The input is identified scene information and audio feature data, and the output is the generated acoustic information.

[0789] Step 4:

[0790] The server creates haptic signals based on the generated acoustic information. The haptic signals are data representations of vibration patterns designed to match the acoustic information. In this step, the generated acoustic information is used as input, and haptic signal data is obtained as output.

[0791] Step 5:

[0792] The server collects real-time emotional data from users and performs emotional analysis. This includes analyzing camera footage and audio input to evaluate the user's emotional state. The input is real-time data from the user, and the output is emotional state information obtained through analysis.

[0793] Step 6:

[0794] The server dynamically adjusts haptic signals based on the user's emotional state information. The vibration intensity and pattern of the haptic signals are modified according to the analysis results. The input consists of emotional state information and haptic signal data, while the output is the adjusted haptic signal.

[0795] Step 7:

[0796] The server packages the tuned haptic signals and audiovisual information as composite data and transmits it to the terminal. A low-latency communication protocol is used to transmit the composite data. The input is the tuned haptic signals and audiovisual information, and the output is the data to be transmitted to the terminal.

[0797] Step 8:

[0798] The terminal decomposes the received composite information and reproduces the visual and auditory information for the user. Simultaneously, it transmits haptic signals to a vibration device, providing the user with physical feedback. The input is data transmitted from the server, and the output is video, audio, and haptic feedback.

[0799] (Application Example 2)

[0800] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0801] Modern content delivery services demand immersive experiences that integrate not only audiovisual but also tactile senses. However, current technology makes it difficult to provide dynamic haptic feedback in real time that responds to the user's emotions. Therefore, there are technical challenges to further enhance the viewer's sensory experience.

[0802] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0803] In this invention, the server includes means for using generation technology to analyze audiovisual signals and generate representative acoustic information, means for creating tactile information based on the generated acoustic information, and data processing means for dynamically adjusting the tactile information based on the user's emotions and transmitting the composite information to a terminal device. This makes it possible to provide the user with real-time tactile feedback synchronized with audiovisual data and optimize the perceptual experience.

[0804] "Audiovisual signals" are data that includes image and sound information, and are fundamental elements for users to experience content through sight and hearing.

[0805] "Acoustic information" is data that represents the characteristics of sound, generated by analyzing audiovisual signals, and extracts representative sounds and important sound changes.

[0806] "Haptic information" refers to signals generated to provide feedback that directly appeals to the user's senses, reproducing physical sensations such as vibrations.

[0807] "User emotions" refer to information that indicates the viewer's current feelings and mental state, and are analyzed in real time from facial expressions and biosignals.

[0808] "Composite information" refers to a data package that integrates audiovisual signals, acoustic information, and haptic information, and is transmitted to terminal devices to provide users with an immersive experience.

[0809] "Terminal device" refers to a device used by a user to experience content, and includes smartphones, head-mounted displays, and other similar devices.

[0810] "Data processing" refers to a series of computational processes that analyze, integrate, and transmit multiple pieces of information, and is a central function of a system for optimizing the user experience.

[0811] The system of the present invention aims to provide users with an immersive experience that integrates sight, hearing, and touch. This system mainly consists of two main components: a server and a terminal device. The server receives audiovisual signals, analyzes important scenes using generation technology, and generates representative acoustic information. Subsequently, it creates tactile information based on this acoustic information and dynamically adjusts the tactile information using an emotion engine that analyzes the user's emotions in real time. This adjusted tactile information is packaged together with the audiovisual signals as composite information and transmitted to the terminal device.

[0812] The terminal device receives this composite information, plays back the audio information, and simultaneously materializes the haptic information using a physical device. This allows users to enhance their content experience not only through sight and hearing, but also through touch. Specific hardware includes smartphones and head-mounted displays, while software such as EmotionEngine and HapticFeedback is used.

[0813] As a concrete example, when a user is watching an emotionally moving film, the server analyzes the viewer's facial expressions, captures subtle emotional changes, and adjusts the haptic information accordingly. Soft vibrations are provided during emotional scenes, while stronger vibrations are generated during action scenes, resulting in an interactive experience.

[0814] Another example of providing prompts to a generative AI model is, "Design a system that analyzes user emotions in real time and provides dynamic haptic feedback according to the video scene." Based on this prompt, the AI ​​model generates an appropriate response, enabling the design of enhanced systems.

[0815] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0816] Step 1:

[0817] The server receives audiovisual signals from content creators. These signals include the video and audio that the user will view. The server uses generation technology to analyze these audiovisual signals and identify important scenes. This analysis allows the server to extract video features and acoustic information and generate representative acoustic information corresponding to the important scenes.

[0818] Step 2:

[0819] The server generates haptic information based on the generated acoustic information. This haptic information is designed to provide the user with a sensory experience in the form of vibration and pressure. This process analyzes the intensity and changes in the acoustic information and designs corresponding vibration patterns. As a result, haptic information that the user actually feels is constructed.

[0820] Step 3:

[0821] The server uses an emotion engine to analyze the user's emotions in real time. It receives the user's facial expression data and biosignals as input and evaluates the current emotional state based on that data. The results of the emotion analysis are used to dynamically adjust haptic information. This makes it possible to design optimal haptic feedback that matches the user's emotions.

[0822] Step 4:

[0823] The server integrates haptic information with audiovisual signals and packages it as composite information. This process seamlessly integrates haptic, acoustic, and visual information to generate a dataset that provides the user with a more complete sensory experience. The generated composite information is then ready to be transmitted to the terminal device.

[0824] Step 5:

[0825] The terminal receives composite information transmitted from the server. It decomposes the composite information, plays the acoustic information through an audio output device, and displays the visual information on the display. It also analyzes tactile information and converts it into signals to control a vibration motor, providing the user with physical tactile feedback.

[0826] Step 6:

[0827] Users enjoy a multi-sensory experience through their devices. By experiencing content that integrates sight, hearing, and touch, they can achieve an unprecedented level of immersion. Through this process, users can feel as if they are actually inside the video.

[0828] This series of processes allows for the dynamic adjustment of haptic information based on the user's emotions, while utilizing audiovisual signals, to provide a personalized entertainment experience.

[0829] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0830] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0831] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0832] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0833] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0834] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0835] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0836] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0837] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0838] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0839] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0840] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0841] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0842] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0843] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0844] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using this memory.

[0845] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0846] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0847] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0848] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0849] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0850] The following is further disclosed regarding the embodiments described above.

[0851] (Claim 1)

[0852] A means of analyzing audiovisual data and using generation technology to generate representative sounds,

[0853] A means for generating tactile signals from a generated representative sound,

[0854] A data processing means for transmitting this composite data, including tactile signals, to a terminal device,

[0855] A system that includes this.

[0856] (Claim 2)

[0857] The system according to claim 1, which decomposes complex data received by a terminal device and realizes haptic signals using a physical device.

[0858] (Claim 3)

[0859] The system according to claim 1, which provides the user with haptic feedback synchronized with audiovisual data to improve the sensory experience.

[0860] "Example 1"

[0861] (Claim 1)

[0862] A means of analyzing audiovisual information and using generation techniques to generate characteristic sounds,

[0863] A means for generating vibration signals based on generated sound,

[0864] Information processing means for transmitting complex information including vibration signals to an information processing device,

[0865] A system that includes this.

[0866] (Claim 2)

[0867] The system according to claim 1, wherein an information processing device decomposes complex information received and a vibration signal is realized by a physical device.

[0868] (Claim 3)

[0869] The system according to claim 1, which provides the user with vibration feedback synchronized with audiovisual information to improve the sensory experience.

[0870] "Application Example 1"

[0871] (Claim 1)

[0872] A means of analyzing audiovisual information and using a generation method to generate characteristic sounds,

[0873] A means for generating tactile signals from generated characteristic sounds,

[0874] Information processing means for transmitting this composite information, including tactile signals, to an information terminal,

[0875] By providing haptic feedback, it is a means of improving the user's sensory experience,

[0876] A system that includes this.

[0877] (Claim 2)

[0878] The system according to claim 1, which decomposes complex information received by an information terminal and realizes tactile signals using a device.

[0879] (Claim 3)

[0880] The system according to claim 1, which provides users with haptic feedback synchronized with audiovisual information to improve their sensory experience.

[0881] "Example 2 of combining an emotion engine"

[0882] (Claim 1)

[0883] A means of using a generation method to analyze audiovisual information, identify a specific scene, and generate sound information,

[0884] A means for creating tactile signals from generated acoustic information,

[0885] An analysis device means for analyzing the user's emotional state and dynamically adjusting haptic signals based on that state,

[0886] Information processing means for transmitting composite information, including generated tactile signals, to a receiving device,

[0887] A system that includes this.

[0888] (Claim 2)

[0889] The system according to claim 1, wherein a receiving device decomposes the complex information it receives and an output device realizes a tactile signal.

[0890] (Claim 3)

[0891] The system according to claim 1, which provides users with haptic feedback synchronized with audiovisual information to improve their perceptual experience.

[0892] "Application example 2 when combining with an emotional engine"

[0893] (Claim 1)

[0894] A means of analyzing audiovisual signals and using generation techniques to generate representative acoustic information,

[0895] A means for creating tactile information based on generated acoustic information,

[0896] A data processing means for dynamically adjusting tactile information based on the user's emotions and transmitting the composite information to a terminal device,

[0897] A system that includes this.

[0898] (Claim 2)

[0899] The system according to claim 1, which analyzes composite information received by a terminal device and materializes tactile information using a physical device.

[0900] (Claim 3)

[0901] The system according to claim 1, which provides humans with haptic feedback harmonized with audiovisual signals and optimizes the perceptual experience. [Explanation of Symbols]

[0902] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of analyzing audiovisual data and using generation technology to generate representative sounds, A means for generating tactile signals from a generated representative sound, A data processing means for transmitting this composite data, including tactile signals, to a terminal device, A system that includes this.

2. The system according to claim 1, wherein a terminal device decomposes complex data received and a tactile signal is realized by a physical device.

3. The system according to claim 1, which provides the user with haptic feedback synchronized with audiovisual data to improve the sensory experience.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A