system
The system simplifies video distribution by compressing and converting data on a server, generating a streaming link, allowing easy viewing on large screens for a diverse user base.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-21
- Publication Date
- 2026-05-07
AI Technical Summary
Existing video distribution systems require complex operations and technical knowledge to deliver video information to remote recipients, making them difficult for elderly or technologically unfamiliar users to use.
A system that compresses video data on a user device, transmits it to a server, converts it into multiple television formats, and generates a streaming link for easy viewing on a large screen, accessible through an internal application on the viewer's device.
Enables easy and efficient video sharing on large screens for a wider range of users, including the elderly and those unfamiliar with technology, without requiring complex operations.
Smart Images

Figure 2026074974000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] There has been a problem that it has been difficult to create an environment in which elderly people and those unfamiliar with technology can easily view videos taken by family members and friends who are far away on a large screen. Usually, technical understanding is required to use a smartphone or a specific streaming device, and the complexity of the operation is a factor. There is a need to solve such problems and provide a system that can be easily used by a wider range of users.
Means for Solving the Problems
[0005] This invention provides a means for automatically compressing video data acquired from a user device and transmitting it to a server device via the internet. The server device converts the received video data into multiple television formats, generates a streaming link based on that data, and notifies the viewer of it. The viewer's video display device reads this notified link using an internal application, thereby streaming the video to the device and providing a means for easy display on a large screen.
[0006] A "user device" is an electronic device used by a user to capture video and manipulate data.
[0007] "Video data" refers to visual media information captured by a user's device, and is material used to allow others to view it.
[0008] "Compression" is the process of reducing the size of video data and optimizing it to make it easier to transmit.
[0009] The "Internet" is a global network used for data transmission and communication.
[0010] A "server device" is a central control unit that receives, stores, and processes video data transmitted from user devices.
[0011] A "television format" is a format that organizes video data according to specific specifications, making it playable on a television.
[0012] "Streaming" is a technology that allows viewers to watch video or audio data while it is being delivered in real time.
[0013] A "link" is reference information used to access specific data or services on the internet.
[0014] A "viewer" is someone who watches the video using the provided link.
[0015] The "video display device" is a device used by viewers to watch videos on a large screen.
[0016] The "internal application" is software incorporated in the video display device and is used to read links and play videos.
Brief Description of Drawings
[0017] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. <00000�4> [Figure 9] It shows an emotion map to which multiple emotions are mapped. [Figure 10] It shows an emotion map to which multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] [[ID=4r]]It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13]It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.
Mode for Carrying Out the Invention
[0018] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0019] First, the language used in the following description will be explained.
[0020] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0021] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0022] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0023] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0025] [First Embodiment]
[0026] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0027] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0028] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0029] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0030] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0032] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0033] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0034] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0035] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0036] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0037] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0038] The system of the present invention consists of a user device, a server device, and a viewer's video display device. The user device functions as a smartphone or camera device and acquires video data of events or travel. The user uses a dedicated application on the terminal to select the captured video data, compress it, and transmit it to the server device via the internet.
[0039] The server device stores the received video data in storage and converts it to a television format. Based on this converted data, it generates a streaming link and notifies the viewer. The viewer uses this link to watch in real time on their home video display device.
[0040] The viewer's video display device functions as a smart TV or streaming device, accessing the notified link through an internal application. This application reads the link, receives streaming data from the server, and displays the video on the TV screen.
[0041] For example, if a user films their grandchild at a school sports day and wants to share the video with grandparents living far away, the user uploads the video using a dedicated app. The server converts the data and sends a generated link to the grandparents' email address. The grandparents can then open the link in a TV application and enjoy watching their grandchild's video on a large screen.
[0042] In this way, the system of the present invention can be easily operated even by viewers unfamiliar with technology, and makes it possible to share what family members are doing on a large screen even if they are far away.
[0043] The following describes the processing flow.
[0044] Step 1:
[0045] Users use their smartphones or camera devices to film events or travels. They then launch a dedicated application and select the filmed video data.
[0046] Step 2:
[0047] The terminal receives the selected video data and compresses it to efficiently transmit it while maintaining data quality. The compressed video data is then sent to the server via the internet.
[0048] Step 3:
[0049] The server stores video data received via the internet in its storage. The server analyzes the received data and performs a process to convert it into multiple television formats (e.g., H.264, MP4).
[0050] Step 4:
[0051] The server generates a streaming link accessible to viewers based on the converted video data. The server then notifies viewers of this link via email or messaging services.
[0052] Step 5:
[0053] Viewers activate their home video display device and check the provided link. They open the internal application on their smart TV or streaming device.
[0054] Step 6:
[0055] The device reads links entered or scanned by the viewer using its internal application. The device receives streaming data from the server and plays the video on the display device's screen in real time.
[0056] (Example 1)
[0057] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0058] Conventional video distribution systems have presented challenges in that they require complex operations and specialized knowledge to efficiently deliver video information captured by users to recipients in remote locations. Therefore, there is a need for a simple and effective video distribution system that is easy to use even for users unfamiliar with technology or elderly recipients.
[0059] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0060] In this invention, the server includes means for compressing video information acquired from a terminal and transmitting it to a data processing device via an information network; means for the data processing device to convert the received video information into multiple visual output formats; and means for generating connection information for distribution based on the converted video information and notifying the recipient. This makes it possible for users to easily operate and distribute videos in real time to family and friends in remote locations without technical constraints.
[0061] A "terminal" refers to an information processing device used by users to acquire video information during events or trips, compress it, and send it to a data processing device.
[0062] "Video information" refers to visual recording data acquired by a device, which is compressed and then distributed.
[0063] An "information network" refers to a means of communication for transmitting video information from a terminal to a data processing device, and includes a wide range of networks such as the internet.
[0064] A "data processing device" refers to a device that receives video information transmitted from a terminal, converts it into multiple visual output formats, and generates connection information for distribution.
[0065] A "visual output format" is a data format used when video information is distributed and visualized, enabling proper display on different devices.
[0066] "Distribution connection information" refers to the links or metadata necessary to deliver the converted video information to recipients, and is generated to make it accessible to recipients.
[0067] A "receiver" refers to a person or device that receives content by visualizing video information transmitted from a device.
[0068] This invention relates to a system for easily distributing video information to recipients in remote locations. Users acquire video information using smartphones or digital cameras as terminals. These terminals are equipped with dedicated applications for efficiently compressing the acquired video information. Standard algorithms such as H.264 and HEVC are used for compression, and after compression, the data is transmitted to a data processing device via an information network.
[0069] The data processing unit functions as a server, saving received video information to a storage system. After saving, software such as FFmpeg is used to convert the video information into multiple visual output formats. This ensures that the information can be properly visualized on any recipient's display device. The converted video information is then used to generate connection information for distribution, which is then notified to the recipient via email or messaging services.
[0070] The recipient uses the verified connection information to visualize the video information in real time using a display device equipped with the internal program, such as a smart TV or streaming device. This allows users to share videos with family and friends in different locations and enjoy them on a larger screen.
[0071] For example, if a user films their child's school play and wants to share it with relatives living abroad, they would upload the video to a server via a dedicated app. An example of a prompt message would be, "Please tell me the steps to take to share a video of my child's concert with relatives living abroad." This creates an environment where even recipients unfamiliar with technology can easily visualize the process.
[0072] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0073] Step 1:
[0074] Users acquire video information using their smartphones or digital cameras. The acquired videos are previewed within a dedicated application, and the user selects which videos to send. The recorded video files are used as input, and the selected videos are prepared within the application as output.
[0075] Step 2:
[0076] The device compresses the selected video. Algorithms such as H.264 and HEVC are used for compression. The input is the selected video file, and the output is a compressed file. This compression reduces the video file size and improves transfer efficiency.
[0077] Step 3:
[0078] The user's terminal sends a compressed video file to a data processing device via an information network. The input is the compressed video file, and the output is the completion of the file transfer to the server. Security is ensured by using protocols such as HTTPS during this process.
[0079] Step 4:
[0080] The server saves the received video to its storage. The input is the transmitted video file, and the output is the saved file that exists in storage. Data integrity is verified during this saving process.
[0081] Step 5:
[0082] The server converts the saved video into a visual output format. Software such as FFmpeg is used to convert it to television formats and other streaming formats. The input is the saved video file, and the output consists of multiple converted visual output format files.
[0083] Step 6:
[0084] The server generates connection information for distribution based on the converted video and notifies the recipient. The input includes the converted video data and the recipient's contact information, while the output is the generated streaming link information and confirmation of its transmission. This information is sent to the recipient via email.
[0085] Step 7:
[0086] The recipient's display device streams the video from its internal program using the notified connection information. The input is the transmitted streaming link information, and based on this, it receives real-time streaming data from the server. The output is the video playback on the recipient's display device. The recipient can easily view the video.
[0087] (Application Example 1)
[0088] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0089] Conventional information sharing systems struggled to quickly and efficiently visualize high-quality information data on multiple display devices. Furthermore, image quality degradation and delays during data compression and conversion processes were problematic, highlighting the need for improved user experience. In particular, there was a need to create a user-friendly system for elderly individuals and those unfamiliar with technology.
[0090] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0091] In this invention, the server includes means for efficiently compressing information data acquired from a user terminal and transmitting it to a central device via a data communication network; means for the central device to convert the received information data into multiple display formats; and means for providing dynamically generated access information based on the converted information data and notifying the information requester. This enables even information requesters unfamiliar with the technology to quickly and efficiently visualize high-quality information data.
[0092] A "user terminal" is a device that has the function of acquiring information data, compressing it efficiently, and transmitting it.
[0093] "High-efficiency compression" refers to a process that reduces data size while maintaining data quality, thereby improving communication speed.
[0094] A "data communication network" is a network used to send and receive information data between devices located in remote locations.
[0095] A "central device" is a device that has the function of converting received information data into multiple display formats and dynamically generating access information.
[0096] "Display format" refers to various formats and protocols used to visualize information data.
[0097] "Dynamically generated access information" refers to information that is generated on the spot by the information requester to provide the necessary links and authentication information to visualize the information data.
[0098] An "information requester" is a person or device that visualizes information data using the provided access information.
[0099] A "display device" is a device that has the function of receiving and visualizing information data.
[0100] An "internal processing program" is software that reads access information and performs visualization of the information data.
[0101] The system of the present invention consists of a user terminal, a central unit, and a display device. The user terminal is a device equipped with a camera and communication functions, and is responsible for acquiring information data and compressing it efficiently. Specifically, the user takes videos of daily life or special events with the user terminal, and the size of the video data is reduced using H.264 or H.265 video compression technology. The compressed information data is then transmitted to the central unit via a data communication network.
[0102] The central system is built on cloud platforms such as Amazon Web Services (AWS®) or Microsoft Azure®. The central system converts the received information data into various display formats (e.g., MP4 or HLS). This conversion process prepares the information data for proper visualization on diverse display devices. The central system then creates dynamically generated access information to notify the information requester, including links and authentication credentials for accessing the information data.
[0103] The information requester uses this access information to visualize the information data on their display device. The display device is a smart TV or tablet, which reads the access information provided by an internal processing program and receives and displays the streaming data.
[0104] As a concrete example, consider a scenario where a user records a family trip and wants to share this video with family members who live far away. The user compresses the video on their device and sends it to a central device. The central device performs the necessary conversions and generates a link that allows for smooth visualization on display devices. This link is sent to family members via email or message, allowing them to enjoy the trip in real time on their own displays.
[0105] Examples of prompts to input into a generative AI model:
[0106] "Please explain in detail how to design a system that efficiently compresses video footage taken during family trips and streams it to family members in real time."
[0107] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0108] Step 1:
[0109] The user acquires information data using their device. Specifically, the user captures video with a camera device and acquires that video data within the device. The input is the captured video data, and the output is the raw video data.
[0110] Step 2:
[0111] The terminal efficiently compresses information data. The user terminal compresses video data using H.264 or H.265 codecs to reduce data size. The input is the acquired video data, and the output is the compressed video data. This compression process improves communication speed and enables efficient data transmission.
[0112] Step 3:
[0113] The terminal transmits compressed information data to the central device via the data communication network. The user terminal transmits the compressed video data to the server via the internet. The input is the compressed video data, and the output is the status of completion of data transmission to the server.
[0114] Step 4:
[0115] The server converts the received data into multiple display formats. The central unit converts the received video data into MP4 or HLS format. The input is the compressed data sent to the server, and the output is the converted video data. This ensures compatibility with various display devices.
[0116] Step 5:
[0117] The server dynamically generates access information based on the converted information data and notifies the information requester. The server generates a streaming link based on the converted data and notifies the specified information requester of the link information via email or message. The input is the converted video data, and the output is the generated streaming link.
[0118] Step 6:
[0119] The information requester's display device visualizes the information data using the notified link information. The information requester reads the received link using the display device's internal processing program and streams the video data. The input is the received streaming link, and the output is the video playback on the display device.
[0120] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0121] The system of the present invention consists of a user device, a server device, a viewer's video display device, and an emotion engine. The user device functions as a smartphone or camera device and not only acquires video data of events and trips, but also has an emotion engine that analyzes the user's facial expressions and voice tone to recognize emotions. The user device selects the captured video data, compresses the video and emotion data, and transmits it to the server device via the internet.
[0122] The server device stores the received video data and emotion data in storage and converts the video to multiple television formats (e.g., H.264, MP4). Based on the converted video data, the server generates a streaming link and notifies the viewer along with the emotion data.
[0123] The viewer's video display device functions as a smart TV or streaming device, displaying notified links and sentiment data in real time using an internal application. This allows viewers to understand the emotional state of the person filming while simultaneously watching the video.
[0124] For example, if a user films their grandchild's sports day and wants to share their emotions (for example, joy) with their grandparents, the user uploads the video along with emotion data recognized by an emotion engine to a server via a dedicated application. The server generates a streaming link combining the sports day video and information indicating the joyful emotion, and sends it to the grandparents. The grandparents can open this link on their TV and watch their grandchild's video while also sharing the emotions of the moment.
[0125] This invention allows for not only viewing of images, but also the simultaneous transmission of the photographer's emotions, enabling a deeper experience and communication.
[0126] The following describes the processing flow.
[0127] Step 1:
[0128] The user uses a smartphone or camera device to record video during events or trips. Simultaneously, an emotion engine built into the user's device analyzes the user's facial expressions and tone of voice to recognize their emotions.
[0129] Step 2:
[0130] The device acquires the captured video data and emotion data recognized by the emotion engine. The device compresses the video data and prepares it to be sent to the server with the emotion data added.
[0131] Step 3:
[0132] The device sends compressed video and emotional data to the server via the internet. Once the transmission is complete, it sends instructions to the server to proceed to the next process.
[0133] Step 4:
[0134] The server stores the received video and emotion data in storage. The server analyzes this video data and converts it into a format suitable for television.
[0135] Step 5:
[0136] The server generates a streaming link based on the converted video data. Simultaneously, it stores the sentiment data sent by the user, associating it with the streaming link.
[0137] Step 6:
[0138] The server notifies viewers of the generated streaming link and associated sentiment data. This notification is sent via email or messaging services.
[0139] Step 7:
[0140] Viewers activate their home video display device and check the notified link. Viewers use an internal application on their smart TV or streaming device to enter or select the streaming link.
[0141] Step 8:
[0142] The device receives input from the viewer and receives streaming data from the server. The device plays the video in real time, displaying emotion data that matches the video data on the display screen.
[0143] (Example 2)
[0144] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0145] Traditional digital content distribution systems only deliver digital content such as video and audio, making it difficult to convey the emotions and intentions of the creator to the viewer. There is a growing need for viewers to understand the emotions behind the content and gain a deeper experience.
[0146] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0147] In this invention, the server includes means for analyzing digital content data acquired from a user terminal and recognizing emotional information, means for transmitting the digital content data and emotional information to an information processing device via the Internet, and means for converting the digital content data received by the information processing device into multiple display formats. This makes it possible for viewers to simultaneously grasp the emotional state of the person who filmed the content while viewing it.
[0148] A "user terminal" refers to a device that has the function of acquiring and analyzing digital content data.
[0149] "Digital content data" is a general term that refers to information acquired electronically, such as video and audio.
[0150] "Emotional information" refers to data that analyzes a user's emotional state and expresses it as numerical values or categories.
[0151] An "information processing device" refers to a device that stores, converts, and generates distribution links for received digital content data and emotional information.
[0152] "Display format" refers to the specific form in which digital content data is converted into a format that can be played on the viewer's device.
[0153] "Communication link" refers to an access method created to provide digital content data and emotional information to viewers via the internet.
[0154] "Display device user" refers to an individual or organization that receives and views digital content data using the provided link.
[0155] "Internal program" refers to software embedded in a display device that reads communication links and plays digital content data.
[0156] The system of this invention consists of a user terminal, a server, and a display device. The user uses a user terminal, such as a smartphone or camera device, to acquire event and daily digital content data. The user terminal is equipped with software called an emotion engine, which analyzes and recognizes emotional information from the user's facial expressions and voice.
[0157] The user's terminal compresses digital content data using compression technologies such as the H.264 codec and generates emotional information. This compressed data and emotional information are transmitted to a server via the internet. The server stores the received digital content data in its storage and converts it into various display formats such as MP4 and WebM. This makes it easy to view the content on different devices.
[0158] The server generates a communication link based on the converted digital content data. This link, along with sentiment information, is sent to the user of the display device. Common notification methods include email and social media. Upon receiving the link, the user opens it via a smart TV or streaming device, and the content is played by an internal program, while the sentiment information is displayed as an overlay.
[0159] As a concrete example, consider a scenario where a user films a family picnic and wants to share the joyful moment with relatives. The user sends the video along with identified emotion information, specifically "joy," to a server. The server generates a communication link based on the picnic video data and emotion information and notifies the relatives. The relatives open this link on their television and share the joy of that moment while watching the picnic video.
[0160] An example of a prompt for a generative AI model would be: "Describe in detail how a user can share a moment they experienced with others through digital content data and visually convey emotional information."
[0161] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0162] Step 1:
[0163] Users acquire digital content data using smartphones or camera devices. The input consists of video and audio from events and daily life, which are analyzed by an emotion engine. The device analyzes the user's facial expressions and tone of voice to recognize emotional information. The output of this step is raw digital content data and its associated emotional information.
[0164] Step 2:
[0165] The device compresses the acquired digital content data using the H.264 codec. The input is the video and audio data obtained in step 1. This data is compressed to reduce its size, and the compressed digital content data is generated as output. In addition, emotional information is converted to JSON format.
[0166] Step 3:
[0167] The terminal sends compressed digital content data and sentiment information in JSON format to the server via the internet. The input for this step is the compressed data and sentiment information, which are sent to the server for use in subsequent processing. The output is the data stored in a format usable by the server.
[0168] Step 4:
[0169] The server stores the received digital content data and sentiment information in storage. The input consists of compressed data and sentiment information sent from the terminal. The server converts the digital content data into multiple display formats (e.g., MP4, WebM). The output is digital content data that can be viewed on different devices.
[0170] Step 5:
[0171] The server generates a communication link from the converted digital content data and notifies the user of this link and sentiment information. The input is the converted digital content data and sentiment information. The server generates a link based on this and notifies the user via email, social media, etc. The output is the notification of the communication link and sentiment information.
[0172] Step 6:
[0173] The user of the display device receives the notified link and opens it using an internal program on their smart TV or streaming device. This plays the digital content data and displays emotional information visually. The input is the communication link, and the output is the viewable digital content data and the display of emotional information.
[0174] (Application Example 2)
[0175] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0176] In today's content sharing environment, there is a lack of means to provide a richer experience that not only allows viewers to watch videos but also simultaneously conveys the emotions of the person who filmed them. Furthermore, there is a need for a simpler and faster way to achieve this.
[0177] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0178] In this invention, the server includes means for transmitting video image data and emotion data acquired from a user terminal via the Internet, means for converting the received video image data into multiple broadcast formats, and means for generating a streaming identifier based on the converted video image data and emotion data and notifying the recipient. This allows viewers to perceive the emotions of the person filming the video simultaneously with the video itself, enabling them to enjoy a more interactive experience.
[0179] A "user terminal" is an electronic device that an individual can carry and use to record and transmit video and audio.
[0180] "Motion image data" refers to data that records visual image information in digital format.
[0181] "Compression" is a process that reduces data size to enable more efficient transmission or storage.
[0182] "Emotional data" refers to digital information that indicates a user's psychological state, obtained by analyzing their facial expressions and tone of voice.
[0183] The "Internet" is a communication framework that connects networks around the world, enabling the transmission and reception of information.
[0184] A "server device" is a computing device connected to a network for receiving, processing, and transmitting large amounts of data.
[0185] A "broadcast format" is a standard digital format for efficiently distributing moving images and videos.
[0186] A "streaming identifier" is information such as a link or hash that identifies specific video data online and facilitates its distribution.
[0187] A "recipient" is an end-user or device that receives data from a sender.
[0188] A "video output device" is a display device that reproduces received video and audio data visually and audibly.
[0189] To implement this invention, a system is constructed that includes a user terminal, a server device, and a video output device for the receiver.
[0190] Program processing
[0191] 1. User terminal operation
[0192] The user terminal is a portable device such as a smartphone or smart glasses, used to acquire video and image data and user emotion data. This terminal is equipped with a camera and microphone, and simultaneously records video and audio while having emotion analysis software installed (e.g., TENSORFLOW®, Amazon Rekognition). This allows the user's facial expressions and voice tone to be analyzed along with the video, generating emotion data.
[0193] 2. Role of the Server Device
[0194] The server device receives video and emotion data transmitted from user terminals and processes them into an efficient format. Specifically, it converts the data into multiple broadcast formats (e.g., H.264, MP4) and generates a streaming identifier. This allows recipients to easily access the data through the identifier.
[0195] 3. Receiver's video output device
[0196] The recipient's video output device is a smart TV or streaming player, which streams video data and sentiment data using an identifier transmitted from the server. The device has an internal program that analyzes the received identifier and displays the video and sentiment information in real time.
[0197] Specific example
[0198] For example, suppose a user wants to record their performance at a festival and share their joy and surprise with friends. In this case, the user's device records the user's emotions along with the video and uploads the data to a server. The server converts the data to a predetermined format and generates an identifier for sending to friends. By using that identifier to play the video on a smart TV, the friend can share the festival footage along with the user's joy and surprise.
[0199] Example of a prompt
[0200] "A student is recording a science experiment show at a science museum. Capture his smile and surprised expression, and suggest the best way for him to share his excitement with his friends."
[0201] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0202] Step 1:
[0203] The user uses a device to record video and audio data. The device uses its camera and microphone to acquire video and audio as input data. From this input data, emotion analysis software installed on the device analyzes the user's facial expressions and voice tone to generate emotion data. The emotion data is output as a numerical value or category indicating the user's psychological state (e.g., joy, surprise).
[0204] Step 2:
[0205] The terminal compresses the generated video and emotion data. A codec (e.g., H.264) is used for compression to reduce data size and facilitate transmission. At this stage, the input is uncompressed video and emotion data, and the output is a compressed data file.
[0206] Step 3:
[0207] The terminal transmits compressed video data and emotion data to the server over the internet. The input is the data compressed in step 2, and the output is the reception of the data on the server. This transmission is performed using a secure and fast protocol (e.g., HTTPS).
[0208] Step 4:
[0209] The server converts the received video data into multiple broadcast formats. The input data is the compressed data received in step 3, and the output is the data converted to a different format (e.g., MP4, WebM). A media conversion library (e.g., FFmpeg) is used for this process.
[0210] Step 5:
[0211] The server generates a streaming identifier based on the converted video and sentiment data and notifies the recipient. The input is the converted data and sentiment data, and the output is provided as a streaming link or code accessible to the recipient. A unique hash or URL is used to generate the identifier.
[0212] Step 6:
[0213] The recipient receives the notified identifier and starts streaming playback on the video output device. The device takes the streaming link as input and displays the video data and emotion data in real time. The output is the video and emotion information that the recipient can see. Through this process, content including emotions is delivered from the user to the recipient.
[0214] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0215] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0216] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0217] [Second Embodiment]
[0218] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0219] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0220] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0221] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0222] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0223] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0224] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0225] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0226] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0227] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0228] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0229] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0230] The system of the present invention consists of a user device, a server device, and a viewer's video display device. The user device functions as a smartphone or camera device and acquires video data of events or travel. The user uses a dedicated application on the terminal to select the captured video data, compress it, and transmit it to the server device via the internet.
[0231] The server device stores the received video data in storage and converts it to a television format. Based on this converted data, it generates a streaming link and notifies the viewer. The viewer uses this link to watch in real time on their home video display device.
[0232] The viewer's video display device functions as a smart TV or streaming device, accessing the notified link through an internal application. This application reads the link, receives streaming data from the server, and displays the video on the TV screen.
[0233] For example, if a user films their grandchild at a school sports day and wants to share the video with grandparents living far away, the user uploads the video using a dedicated app. The server converts the data and sends a generated link to the grandparents' email address. The grandparents can then open the link in a TV application and enjoy watching their grandchild's video on a large screen.
[0234] In this way, the system of the present invention can be easily operated even by viewers unfamiliar with technology, and makes it possible to share what family members are doing on a large screen even if they are far away.
[0235] The following describes the processing flow.
[0236] Step 1:
[0237] Users use their smartphones or camera devices to film events or travels. They then launch a dedicated application and select the filmed video data.
[0238] Step 2:
[0239] The terminal receives the selected video data and compresses it to efficiently transmit it while maintaining data quality. The compressed video data is then sent to the server via the internet.
[0240] Step 3:
[0241] The server stores video data received via the internet in its storage. The server analyzes the received data and performs a process to convert it into multiple television formats (e.g., H.264, MP4).
[0242] Step 4:
[0243] The server generates a streaming link accessible to viewers based on the converted video data. The server then notifies viewers of this link via email or messaging services.
[0244] Step 5:
[0245] Viewers activate their home video display device and check the provided link. They open the internal application on their smart TV or streaming device.
[0246] Step 6:
[0247] The device reads links entered or scanned by the viewer using its internal application. The device receives streaming data from the server and plays the video on the display device's screen in real time.
[0248] (Example 1)
[0249] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0250] Conventional video distribution systems have presented challenges in that they require complex operations and specialized knowledge to efficiently deliver video information captured by users to recipients in remote locations. Therefore, there is a need for a simple and effective video distribution system that is easy to use even for users unfamiliar with technology or elderly recipients.
[0251] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0252] In this invention, the server includes means for compressing video information acquired from a terminal and transmitting it to a data processing device via an information network; means for the data processing device to convert the received video information into multiple visual output formats; and means for generating connection information for distribution based on the converted video information and notifying the recipient. This makes it possible for users to easily operate and distribute videos in real time to family and friends in remote locations without technical constraints.
[0253] A "terminal" refers to an information processing device used by users to acquire video information during events or trips, compress it, and send it to a data processing device.
[0254] "Video information" refers to visual recording data acquired by a device, which is compressed and then distributed.
[0255] An "information network" refers to a means of communication for transmitting video information from a terminal to a data processing device, and includes a wide range of networks such as the internet.
[0256] A "data processing device" refers to a device that receives video information transmitted from a terminal, converts it into multiple visual output formats, and generates connection information for distribution.
[0257] A "visual output format" is a data format used when video information is distributed and visualized, enabling proper display on different devices.
[0258] "Distribution connection information" refers to the links or metadata necessary to deliver the converted video information to recipients, and is generated to make it accessible to recipients.
[0259] A "receiver" refers to a person or device that receives content by visualizing video information transmitted from a device.
[0260] This invention relates to a system for easily distributing video information to recipients in remote locations. Users acquire video information using smartphones or digital cameras as terminals. These terminals are equipped with dedicated applications for efficiently compressing the acquired video information. Standard algorithms such as H.264 and HEVC are used for compression, and after compression, the data is transmitted to a data processing device via an information network.
[0261] The data processing unit functions as a server, saving received video information to a storage system. After saving, software such as FFmpeg is used to convert the video information into multiple visual output formats. This ensures that the information can be properly visualized on any recipient's display device. The converted video information is then used to generate connection information for distribution, which is then notified to the recipient via email or messaging services.
[0262] The recipient uses the verified connection information to visualize the video information in real time using a display device equipped with the internal program, such as a smart TV or streaming device. This allows users to share videos with family and friends in different locations and enjoy them on a larger screen.
[0263] For example, if a user films their child's school play and wants to share it with relatives living abroad, they would upload the video to a server via a dedicated app. An example of a prompt message would be, "Please tell me the steps to take to share a video of my child's concert with relatives living abroad." This creates an environment where even recipients unfamiliar with technology can easily visualize the process.
[0264] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0265] Step 1:
[0266] Users acquire video information using their smartphones or digital cameras. The acquired videos are previewed within a dedicated application, and the user selects which videos to send. The recorded video files are used as input, and the selected videos are prepared within the application as output.
[0267] Step 2:
[0268] The device compresses the selected video. Algorithms such as H.264 and HEVC are used for compression. The input is the selected video file, and the output is a compressed file. This compression reduces the video file size and improves transfer efficiency.
[0269] Step 3:
[0270] The user's terminal sends a compressed video file to a data processing device via an information network. The input is the compressed video file, and the output is the completion of the file transfer to the server. Security is ensured by using protocols such as HTTPS during this process.
[0271] Step 4:
[0272] The server saves the received video to its storage. The input is the transmitted video file, and the output is the saved file that exists in storage. Data integrity is verified during this saving process.
[0273] Step 5:
[0274] The server converts the saved video into a visual output format. Software such as FFmpeg is used to convert it to television formats and other streaming formats. The input is the saved video file, and the output consists of multiple converted visual output format files.
[0275] Step 6:
[0276] The server generates connection information for distribution based on the converted video and notifies the recipient. The input includes the converted video data and the recipient's contact information, while the output is the generated streaming link information and confirmation of its transmission. This information is sent to the recipient via email.
[0277] Step 7:
[0278] The recipient's display device streams the video from its internal program using the notified connection information. The input is the transmitted streaming link information, and based on this, it receives real-time streaming data from the server. The output is the video playback on the recipient's display device. The recipient can easily view the video.
[0279] (Application Example 1)
[0280] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0281] Conventional information sharing systems struggled to quickly and efficiently visualize high-quality information data on multiple display devices. Furthermore, image quality degradation and delays during data compression and conversion processes were problematic, highlighting the need for improved user experience. In particular, there was a need to create a user-friendly system for elderly individuals and those unfamiliar with technology.
[0282] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0283] In this invention, the server includes means for highly efficiently compressing information data acquired from a user terminal and transmitting it to a central device via a data communication network, means for converting the information data received by the central device into a plurality of display formats, and means for providing access information dynamically generated based on the converted information data and notifying an information requester. Thereby, it becomes possible to quickly and efficiently visualize high-quality information data even for information requesters who are not familiar with the technology.
[0284] The "user terminal" is a device having a function of acquiring information data, highly efficiently compressing it, and transmitting it.
[0285] "Highly efficiently compressing" is a process of reducing the capacity while maintaining the quality of the data and improving the communication speed.
[0286] The "data communication network" is a network for transmitting and receiving information data between devices located in remote locations.
[0287] The "central device" is a device having a function of converting the received information data into a plurality of display formats and dynamically generating access information.
[0288] The "display format" refers to various formats and protocols for visualizing information data.
[0289] The "dynamically generated access information" is something that generates on-the-spot the links and authentication information necessary for an information requester to visualize information data.
[0290] The "information requester" is a person or device in a position to visualize information data using the provided access information.
[0291] The "display device" is a device having a function of receiving and visualizing information data.
[0292] The "internal processing program" is software for reading access information and executing the visualization of information data.
[0293] The system of the present invention consists of a user terminal, a central unit, and a display device. The user terminal is a device equipped with a camera and communication functions, and is responsible for acquiring information data and compressing it efficiently. Specifically, the user takes videos of daily life or special events with the user terminal, and the size of the video data is reduced using H.264 or H.265 video compression technology. The compressed information data is then transmitted to the central unit via a data communication network.
[0294] The central system is built on cloud platforms such as Amazon Web Services (AWS) or Microsoft Azure. The central system converts the received information data into various display formats (e.g., MP4 or HLS). This conversion process ensures the information data is formatted to be properly visualized on diverse display devices. The central system then creates dynamically generated access information to notify the information requester, including links and authentication credentials for accessing the information data.
[0295] The information requester uses this access information to visualize the information data on their display device. The display device is a smart TV or tablet, which reads the access information provided by an internal processing program and receives and displays the streaming data.
[0296] As a concrete example, consider a scenario where a user records a family trip and wants to share this video with family members who live far away. The user compresses the video on their device and sends it to a central device. The central device performs the necessary conversions and generates a link that allows for smooth visualization on display devices. This link is sent to family members via email or message, allowing them to enjoy the trip in real time on their own displays.
[0297] Examples of prompts to input into a generative AI model:
[0298] "Please explain in detail how to design a system that efficiently compresses videos taken during a family trip and streams them to family members in real-time."
[0299] The flow of the specific process in Application Example 1 will be described using FIG. 12.
[0300] Step 1:
[0301] The user acquires information data using the user terminal. Specifically, the user takes a video with a camera device and acquires the video data within the terminal. The input is the captured video data, and the output is the same video data.
[0302] Step 2:
[0303] The terminal efficiently compresses the information data. The user terminal compresses the video data using H.264 or H.265 codec to reduce the data size. The input is the acquired video data, and the output is the compressed video data. This compression process improves the communication speed and enables efficient data transmission.
[0304] Step 3:
[0305] The terminal transmits the compressed information data to the central device via the data communication network. The user terminal transmits the compressed video data to the server via the Internet. The input is the compressed video data, and the output is the data transmission completion status to the server.
[0306] Step 4:
[0307] The server converts the received information data into multiple display formats. The central device converts the received video data into MP4 or HLS format. The input is the compressed data transmitted to the server, and the output is the converted video data. This ensures compatibility with various display devices.
[0308] Step 5:
[0309] The server dynamically generates access information based on the converted information data and notifies the information requester. The server generates a streaming link based on the converted data and notifies the specified information requester of the link information via email or message. The input is the converted video data, and the output is the generated streaming link.
[0310] Step 6:
[0311] The information requester's display device visualizes the information data using the notified link information. The information requester reads the received link using the display device's internal processing program and streams the video data. The input is the received streaming link, and the output is the video playback on the display device.
[0312] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0313] The system of the present invention consists of a user device, a server device, a viewer's video display device, and an emotion engine. The user device functions as a smartphone or camera device and not only acquires video data of events and trips, but also has an emotion engine that analyzes the user's facial expressions and voice tone to recognize emotions. The user device selects the captured video data, compresses the video and emotion data, and transmits it to the server device via the internet.
[0314] The server device stores the received video data and emotion data in storage and converts the video to multiple television formats (e.g., H.264, MP4). Based on the converted video data, the server generates a streaming link and notifies the viewer along with the emotion data.
[0315] The viewer's video display device functions as a smart TV or streaming device, displaying notified links and sentiment data in real time using an internal application. This allows viewers to understand the emotional state of the person filming while simultaneously watching the video.
[0316] For example, if a user films their grandchild's sports day and wants to share their emotions (for example, joy) with their grandparents, the user uploads the video along with emotion data recognized by an emotion engine to a server via a dedicated application. The server generates a streaming link combining the sports day video and information indicating the joyful emotion, and sends it to the grandparents. The grandparents can open this link on their TV and watch their grandchild's video while also sharing the emotions of the moment.
[0317] This invention allows for not only viewing of images, but also the simultaneous transmission of the photographer's emotions, enabling a deeper experience and communication.
[0318] The following describes the processing flow.
[0319] Step 1:
[0320] The user uses a smartphone or camera device to record video during events or trips. Simultaneously, an emotion engine built into the user's device analyzes the user's facial expressions and tone of voice to recognize their emotions.
[0321] Step 2:
[0322] The device acquires the captured video data and emotion data recognized by the emotion engine. The device compresses the video data and prepares it to be sent to the server with the emotion data added.
[0323] Step 3:
[0324] The device sends compressed video and emotional data to the server via the internet. Once the transmission is complete, it sends instructions to the server to proceed to the next process.
[0325] Step 4:
[0326] The server stores the received video and emotion data in storage. The server analyzes this video data and converts it into a format suitable for television.
[0327] Step 5:
[0328] The server generates a streaming link based on the converted video data. Simultaneously, it stores the sentiment data sent by the user, associating it with the streaming link.
[0329] Step 6:
[0330] The server notifies viewers of the generated streaming link and associated sentiment data. This notification is sent via email or messaging services.
[0331] Step 7:
[0332] Viewers activate their home video display device and check the notified link. Viewers use an internal application on their smart TV or streaming device to enter or select the streaming link.
[0333] Step 8:
[0334] The device receives input from the viewer and receives streaming data from the server. The device plays the video in real time, displaying emotion data that matches the video data on the display screen.
[0335] (Example 2)
[0336] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0337] Traditional digital content distribution systems only deliver digital content such as video and audio, making it difficult to convey the emotions and intentions of the creator to the viewer. There is a growing need for viewers to understand the emotions behind the content and gain a deeper experience.
[0338] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0339] In this invention, the server includes means for analyzing digital content data acquired from a user terminal and recognizing emotional information, means for transmitting the digital content data and emotional information to an information processing device via the Internet, and means for converting the digital content data received by the information processing device into multiple display formats. This makes it possible for viewers to simultaneously grasp the emotional state of the person who filmed the content while viewing it.
[0340] A "user terminal" refers to a device that has the function of acquiring and analyzing digital content data.
[0341] "Digital content data" is a general term that refers to information acquired electronically, such as video and audio.
[0342] "Emotional information" refers to data that analyzes a user's emotional state and expresses it as numerical values or categories.
[0343] An "information processing device" refers to a device that stores, converts, and generates distribution links for received digital content data and emotional information.
[0344] "Display format" refers to the specific form in which digital content data is converted into a format that can be played on the viewer's device.
[0345] "Communication link" refers to an access method created to provide digital content data and emotional information to viewers via the internet.
[0346] "Display device user" refers to an individual or organization that receives and views digital content data using the provided link.
[0347] "Internal program" refers to software embedded in a display device that reads communication links and plays digital content data.
[0348] The system of this invention consists of a user terminal, a server, and a display device. The user uses a user terminal, such as a smartphone or camera device, to acquire event and daily digital content data. The user terminal is equipped with software called an emotion engine, which analyzes and recognizes emotional information from the user's facial expressions and voice.
[0349] The user's terminal compresses digital content data using compression technologies such as the H.264 codec and generates emotional information. This compressed data and emotional information are transmitted to a server via the internet. The server stores the received digital content data in its storage and converts it into various display formats such as MP4 and WebM. This makes it easy to view the content on different devices.
[0350] The server generates a communication link based on the converted digital content data. This link, along with sentiment information, is sent to the user of the display device. Common notification methods include email and social media. Upon receiving the link, the user opens it via a smart TV or streaming device, and the content is played by an internal program, while the sentiment information is displayed as an overlay.
[0351] As a concrete example, consider a scenario where a user films a family picnic and wants to share the joyful moment with relatives. The user sends the video along with identified emotion information, specifically "joy," to a server. The server generates a communication link based on the picnic video data and emotion information and notifies the relatives. The relatives open this link on their television and share the joy of that moment while watching the picnic video.
[0352] An example of a prompt for a generative AI model would be: "Describe in detail how a user can share a moment they experienced with others through digital content data and visually convey emotional information."
[0353] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0354] Step 1:
[0355] Users acquire digital content data using smartphones or camera devices. The input consists of video and audio from events and daily life, which are analyzed by an emotion engine. The device analyzes the user's facial expressions and tone of voice to recognize emotional information. The output of this step is raw digital content data and its associated emotional information.
[0356] Step 2:
[0357] The device compresses the acquired digital content data using the H.264 codec. The input is the video and audio data obtained in step 1. This data is compressed to reduce its size, and the compressed digital content data is generated as output. In addition, emotional information is converted to JSON format.
[0358] Step 3:
[0359] The terminal sends compressed digital content data and sentiment information in JSON format to the server via the internet. The input for this step is the compressed data and sentiment information, which are sent to the server for use in subsequent processing. The output is the data stored in a format usable by the server.
[0360] Step 4:
[0361] The server stores the received digital content data and sentiment information in storage. The input consists of compressed data and sentiment information sent from the terminal. The server converts the digital content data into multiple display formats (e.g., MP4, WebM). The output is digital content data that can be viewed on different devices.
[0362] Step 5:
[0363] The server generates a communication link from the converted digital content data and notifies the user of this link and sentiment information. The input is the converted digital content data and sentiment information. The server generates a link based on this and notifies the user via email, social media, etc. The output is the notification of the communication link and sentiment information.
[0364] Step 6:
[0365] The user of the display device receives the notified link and opens it using an internal program on their smart TV or streaming device. This plays the digital content data and displays emotional information visually. The input is the communication link, and the output is the viewable digital content data and the display of emotional information.
[0366] (Application Example 2)
[0367] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0368] In today's content sharing environment, there is a lack of means to provide a richer experience that not only allows viewers to watch videos but also simultaneously conveys the emotions of the person who filmed them. Furthermore, there is a need for a simpler and faster way to achieve this.
[0369] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0370] In this invention, the server includes means for transmitting video image data and emotion data acquired from a user terminal via the Internet, means for converting the received video image data into multiple broadcast formats, and means for generating a streaming identifier based on the converted video image data and emotion data and notifying the recipient. This allows viewers to perceive the emotions of the person filming the video simultaneously with the video itself, enabling them to enjoy a more interactive experience.
[0371] A "user terminal" is an electronic device that an individual can carry and use to record and transmit video and audio.
[0372] "Motion image data" refers to data that records visual image information in digital format.
[0373] "Compression" is a process that reduces data size to enable more efficient transmission or storage.
[0374] "Emotional data" refers to digital information that indicates a user's psychological state, obtained by analyzing their facial expressions and tone of voice.
[0375] The "Internet" is a communication framework that connects networks around the world, enabling the transmission and reception of information.
[0376] A "server device" is a computing device connected to a network for receiving, processing, and transmitting large amounts of data.
[0377] A "broadcast format" is a standard digital format for efficiently distributing moving images and videos.
[0378] A "streaming identifier" is information such as a link or hash that identifies specific video data online and facilitates its distribution.
[0379] A "recipient" is an end-user or device that receives data from a sender.
[0380] A "video output device" is a display device that reproduces received video and audio data visually and audibly.
[0381] To implement this invention, a system is constructed that includes a user terminal, a server device, and a video output device for the receiver.
[0382] Program processing
[0383] 1. User terminal operation
[0384] The user terminal is a portable device such as a smartphone or smart glasses, used to acquire video and image data and user emotion data. This terminal is equipped with a camera and microphone, and simultaneously records video and audio, while also having emotion analysis software (e.g., TensorFlow, Amazon Rekognition) installed. This allows the user's facial expressions and voice tone to be analyzed along with the video, generating emotion data.
[0385] 2. Role of the Server Device
[0386] The server device receives video and emotion data transmitted from user terminals and processes them into an efficient format. Specifically, it converts the data into multiple broadcast formats (e.g., H.264, MP4) and generates a streaming identifier. This allows recipients to easily access the data through the identifier.
[0387] 3. Receiver's video output device
[0388] The recipient's video output device is a smart TV or streaming player, which streams video data and sentiment data using an identifier transmitted from the server. The device has an internal program that analyzes the received identifier and displays the video and sentiment information in real time.
[0389] Specific example
[0390] For example, suppose a user wants to record their performance at a festival and share their joy and surprise with friends. In this case, the user's device records the user's emotions along with the video and uploads the data to a server. The server converts the data to a predetermined format and generates an identifier for sending to friends. By using that identifier to play the video on a smart TV, the friend can share the festival footage along with the user's joy and surprise.
[0391] Example of a prompt
[0392] "A student is recording a science experiment show at a science museum. Capture his smile and surprised expression, and suggest the best way for him to share his excitement with his friends."
[0393] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0394] Step 1:
[0395] The user uses a device to record video and audio data. The device uses its camera and microphone to acquire video and audio as input data. From this input data, emotion analysis software installed on the device analyzes the user's facial expressions and voice tone to generate emotion data. The emotion data is output as a numerical value or category indicating the user's psychological state (e.g., joy, surprise).
[0396] Step 2:
[0397] The terminal compresses the generated video and emotion data. A codec (e.g., H.264) is used for compression to reduce data size and facilitate transmission. At this stage, the input is uncompressed video and emotion data, and the output is a compressed data file.
[0398] Step 3:
[0399] The terminal transmits compressed video data and emotion data to the server over the internet. The input is the data compressed in step 2, and the output is the reception of the data on the server. This transmission is performed using a secure and fast protocol (e.g., HTTPS).
[0400] Step 4:
[0401] The server converts the received video data into multiple broadcast formats. The input data is the compressed data received in step 3, and the output is the data converted to a different format (e.g., MP4, WebM). A media conversion library (e.g., FFmpeg) is used for this process.
[0402] Step 5:
[0403] The server generates a streaming identifier based on the converted video and sentiment data and notifies the recipient. The input is the converted data and sentiment data, and the output is provided as a streaming link or code accessible to the recipient. A unique hash or URL is used to generate the identifier.
[0404] Step 6:
[0405] The recipient receives the notified identifier and starts streaming playback on the video output device. The device takes the streaming link as input and displays the video data and emotion data in real time. The output is the video and emotion information that the recipient can see. Through this process, content including emotions is delivered from the user to the recipient.
[0406] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0407] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0408] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0409] [Third Embodiment]
[0410] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0411] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0412] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0413] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0414] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0415] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0416] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0417] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0418] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0419] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0420] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0421] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0422] The system of the present invention consists of a user device, a server device, and a viewer's video display device. The user device functions as a smartphone or camera device and acquires video data of events or travel. The user uses a dedicated application on the terminal to select the captured video data, compress it, and transmit it to the server device via the internet.
[0423] The server device stores the received video data in storage and converts it to a television format. Based on this converted data, it generates a streaming link and notifies the viewer. The viewer uses this link to watch in real time on their home video display device.
[0424] The viewer's video display device functions as a smart TV or streaming device, accessing the notified link through an internal application. This application reads the link, receives streaming data from the server, and displays the video on the TV screen.
[0425] For example, if a user films their grandchild at a school sports day and wants to share the video with grandparents living far away, the user uploads the video using a dedicated app. The server converts the data and sends a generated link to the grandparents' email address. The grandparents can then open the link in a TV application and enjoy watching their grandchild's video on a large screen.
[0426] In this way, the system of the present invention can be easily operated even by viewers unfamiliar with technology, and makes it possible to share what family members are doing on a large screen even if they are far away.
[0427] The following describes the processing flow.
[0428] Step 1:
[0429] Users use their smartphones or camera devices to film events or travels. They then launch a dedicated application and select the filmed video data.
[0430] Step 2:
[0431] The terminal receives the selected video data and compresses it to efficiently transmit it while maintaining data quality. The compressed video data is then sent to the server via the internet.
[0432] Step 3:
[0433] The server stores video data received via the internet in its storage. The server analyzes the received data and performs a process to convert it into multiple television formats (e.g., H.264, MP4).
[0434] Step 4:
[0435] The server generates a streaming link accessible to viewers based on the converted video data. The server then notifies viewers of this link via email or messaging services.
[0436] Step 5:
[0437] Viewers activate their home video display device and check the provided link. They open the internal application on their smart TV or streaming device.
[0438] Step 6:
[0439] The device reads links entered or scanned by the viewer using its internal application. The device receives streaming data from the server and plays the video on the display device's screen in real time.
[0440] (Example 1)
[0441] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0442] Conventional video distribution systems have presented challenges in that they require complex operations and specialized knowledge to efficiently deliver video information captured by users to recipients in remote locations. Therefore, there is a need for a simple and effective video distribution system that is easy to use even for users unfamiliar with technology or elderly recipients.
[0443] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0444] In this invention, the server includes means for compressing video information acquired from a terminal and transmitting it to a data processing device via an information network; means for the data processing device to convert the received video information into multiple visual output formats; and means for generating connection information for distribution based on the converted video information and notifying the recipient. This makes it possible for users to easily operate and distribute videos in real time to family and friends in remote locations without technical constraints.
[0445] A "terminal" refers to an information processing device used by users to acquire video information during events or trips, compress it, and send it to a data processing device.
[0446] "Video information" refers to visual recording data acquired by a device, which is compressed and then distributed.
[0447] An "information network" refers to a means of communication for transmitting video information from a terminal to a data processing device, and includes a wide range of networks such as the internet.
[0448] A "data processing device" refers to a device that receives video information transmitted from a terminal, converts it into multiple visual output formats, and generates connection information for distribution.
[0449] A "visual output format" is a data format used when video information is distributed and visualized, enabling proper display on different devices.
[0450] "Distribution connection information" refers to the links or metadata necessary to deliver the converted video information to recipients, and is generated to make it accessible to recipients.
[0451] A "receiver" refers to a person or device that receives content by visualizing video information transmitted from a device.
[0452] This invention relates to a system for easily distributing video information to recipients in remote locations. Users acquire video information using smartphones or digital cameras as terminals. These terminals are equipped with dedicated applications for efficiently compressing the acquired video information. Standard algorithms such as H.264 and HEVC are used for compression, and after compression, the data is transmitted to a data processing device via an information network.
[0453] The data processing unit functions as a server, saving received video information to a storage system. After saving, software such as FFmpeg is used to convert the video information into multiple visual output formats. This ensures that the information can be properly visualized on any recipient's display device. The converted video information is then used to generate connection information for distribution, which is then notified to the recipient via email or messaging services.
[0454] The recipient uses the verified connection information to visualize the video information in real time using a display device equipped with the internal program, such as a smart TV or streaming device. This allows users to share videos with family and friends in different locations and enjoy them on a larger screen.
[0455] For example, if a user films their child's school play and wants to share it with relatives living abroad, they would upload the video to a server via a dedicated app. An example of a prompt message would be, "Please tell me the steps to take to share a video of my child's concert with relatives living abroad." This creates an environment where even recipients unfamiliar with technology can easily visualize the process.
[0456] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0457] Step 1:
[0458] Users acquire video information using their smartphones or digital cameras. The acquired videos are previewed within a dedicated application, and the user selects which videos to send. The recorded video files are used as input, and the selected videos are prepared within the application as output.
[0459] Step 2:
[0460] The device compresses the selected video. Algorithms such as H.264 and HEVC are used for compression. The input is the selected video file, and the output is a compressed file. This compression reduces the video file size and improves transfer efficiency.
[0461] Step 3:
[0462] The user's terminal sends a compressed video file to a data processing device via an information network. The input is the compressed video file, and the output is the completion of the file transfer to the server. Security is ensured by using protocols such as HTTPS during this process.
[0463] Step 4:
[0464] The server saves the received video to its storage. The input is the transmitted video file, and the output is the saved file that exists in storage. Data integrity is verified during this saving process.
[0465] Step 5:
[0466] The server converts the saved video into a visual output format. Software such as FFmpeg is used to convert it to television formats and other streaming formats. The input is the saved video file, and the output consists of multiple converted visual output format files.
[0467] Step 6:
[0468] The server generates connection information for distribution based on the converted video and notifies the recipient. The input includes the converted video data and the recipient's contact information, while the output is the generated streaming link information and confirmation of its transmission. This information is sent to the recipient via email.
[0469] Step 7:
[0470] The recipient's display device streams the video from its internal program using the notified connection information. The input is the transmitted streaming link information, and based on this, it receives real-time streaming data from the server. The output is the video playback on the recipient's display device. The recipient can easily view the video.
[0471] (Application Example 1)
[0472] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0473] Conventional information sharing systems struggled to quickly and efficiently visualize high-quality information data on multiple display devices. Furthermore, image quality degradation and delays during data compression and conversion processes were problematic, highlighting the need for improved user experience. In particular, there was a need to create a user-friendly system for elderly individuals and those unfamiliar with technology.
[0474] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0475] In this invention, the server includes means for efficiently compressing information data acquired from a user terminal and transmitting it to a central device via a data communication network; means for the central device to convert the received information data into multiple display formats; and means for providing dynamically generated access information based on the converted information data and notifying the information requester. This enables even information requesters unfamiliar with the technology to quickly and efficiently visualize high-quality information data.
[0476] A "user terminal" is a device that has the function of acquiring information data, compressing it efficiently, and transmitting it.
[0477] "High-efficiency compression" refers to a process that reduces data size while maintaining data quality, thereby improving communication speed.
[0478] A "data communication network" is a network used to send and receive information data between devices located in remote locations.
[0479] A "central device" is a device that has the function of converting received information data into multiple display formats and dynamically generating access information.
[0480] "Display format" refers to various formats and protocols used to visualize information data.
[0481] "Dynamically generated access information" refers to information that is generated on the spot by the information requester to provide the necessary links and authentication information to visualize the information data.
[0482] An "information requester" is a person or device that visualizes information data using the provided access information.
[0483] A "display device" is a device that has the function of receiving and visualizing information data.
[0484] An "internal processing program" is software that reads access information and performs visualization of the information data.
[0485] The system of the present invention consists of a user terminal, a central unit, and a display device. The user terminal is a device equipped with a camera and communication functions, and is responsible for acquiring information data and compressing it efficiently. Specifically, the user takes videos of daily life or special events with the user terminal, and the size of the video data is reduced using H.264 or H.265 video compression technology. The compressed information data is then transmitted to the central unit via a data communication network.
[0486] The central system is built on cloud platforms such as Amazon Web Services (AWS) or Microsoft Azure. The central system converts the received information data into various display formats (e.g., MP4 or HLS). This conversion process ensures the information data is formatted to be properly visualized on diverse display devices. The central system then creates dynamically generated access information to notify the information requester, including links and authentication credentials for accessing the information data.
[0487] The information requester uses this access information to visualize the information data on their display device. The display device is a smart TV or tablet, which reads the access information provided by an internal processing program and receives and displays the streaming data.
[0488] As a concrete example, consider a scenario where a user records a family trip and wants to share this video with family members who live far away. The user compresses the video on their device and sends it to a central device. The central device performs the necessary conversions and generates a link that allows for smooth visualization on display devices. This link is sent to family members via email or message, allowing them to enjoy the trip in real time on their own displays.
[0489] Examples of prompts to input into a generative AI model:
[0490] "Please explain in detail how to design a system that efficiently compresses video footage taken during family trips and streams it to family members in real time."
[0491] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0492] Step 1:
[0493] The user acquires information data using their device. Specifically, the user captures video with a camera device and acquires that video data within the device. The input is the captured video data, and the output is the raw video data.
[0494] Step 2:
[0495] The terminal efficiently compresses information data. The user terminal compresses video data using H.264 or H.265 codecs to reduce data size. The input is the acquired video data, and the output is the compressed video data. This compression process improves communication speed and enables efficient data transmission.
[0496] Step 3:
[0497] The terminal transmits compressed information data to the central device via the data communication network. The user terminal transmits the compressed video data to the server via the internet. The input is the compressed video data, and the output is the status of completion of data transmission to the server.
[0498] Step 4:
[0499] The server converts the received data into multiple display formats. The central unit converts the received video data into MP4 or HLS format. The input is the compressed data sent to the server, and the output is the converted video data. This ensures compatibility with various display devices.
[0500] Step 5:
[0501] The server dynamically generates access information based on the converted information data and notifies the information requester. The server generates a streaming link based on the converted data and notifies the specified information requester of the link information via email or message. The input is the converted video data, and the output is the generated streaming link.
[0502] Step 6:
[0503] The information requester's display device visualizes the information data using the notified link information. The information requester reads the received link using the display device's internal processing program and streams the video data. The input is the received streaming link, and the output is the video playback on the display device.
[0504] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0505] The system of the present invention consists of a user device, a server device, a viewer's video display device, and an emotion engine. The user device functions as a smartphone or camera device and not only acquires video data of events and trips, but also has an emotion engine that analyzes the user's facial expressions and voice tone to recognize emotions. The user device selects the captured video data, compresses the video and emotion data, and transmits it to the server device via the internet.
[0506] The server device stores the received video data and emotion data in storage and converts the video to multiple television formats (e.g., H.264, MP4). Based on the converted video data, the server generates a streaming link and notifies the viewer along with the emotion data.
[0507] The viewer's video display device functions as a smart TV or streaming device, displaying notified links and sentiment data in real time using an internal application. This allows viewers to understand the emotional state of the person filming while simultaneously watching the video.
[0508] For example, if a user films their grandchild's sports day and wants to share their emotions (for example, joy) with their grandparents, the user uploads the video along with emotion data recognized by an emotion engine to a server via a dedicated application. The server generates a streaming link combining the sports day video and information indicating the joyful emotion, and sends it to the grandparents. The grandparents can open this link on their TV and watch their grandchild's video while also sharing the emotions of the moment.
[0509] This invention allows for not only viewing of images, but also the simultaneous transmission of the photographer's emotions, enabling a deeper experience and communication.
[0510] The following describes the processing flow.
[0511] Step 1:
[0512] The user uses a smartphone or camera device to record video during events or trips. Simultaneously, an emotion engine built into the user's device analyzes the user's facial expressions and tone of voice to recognize their emotions.
[0513] Step 2:
[0514] The device acquires the captured video data and emotion data recognized by the emotion engine. The device compresses the video data and prepares it to be sent to the server with the emotion data added.
[0515] Step 3:
[0516] The device sends compressed video and emotional data to the server via the internet. Once the transmission is complete, it sends instructions to the server to proceed to the next process.
[0517] Step 4:
[0518] The server stores the received video and emotion data in storage. The server analyzes this video data and converts it into a format suitable for television.
[0519] Step 5:
[0520] The server generates a streaming link based on the converted video data. Simultaneously, it stores the sentiment data sent by the user, associating it with the streaming link.
[0521] Step 6:
[0522] The server notifies viewers of the generated streaming link and associated sentiment data. This notification is sent via email or messaging services.
[0523] Step 7:
[0524] Viewers activate their home video display device and check the notified link. Viewers use an internal application on their smart TV or streaming device to enter or select the streaming link.
[0525] Step 8:
[0526] The device receives input from the viewer and receives streaming data from the server. The device plays the video in real time, displaying emotion data that matches the video data on the display screen.
[0527] (Example 2)
[0528] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0529] Traditional digital content distribution systems only deliver digital content such as video and audio, making it difficult to convey the emotions and intentions of the creator to the viewer. There is a growing need for viewers to understand the emotions behind the content and gain a deeper experience.
[0530] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0531] In this invention, the server includes means for analyzing digital content data acquired from a user terminal and recognizing emotional information, means for transmitting the digital content data and emotional information to an information processing device via the Internet, and means for converting the digital content data received by the information processing device into multiple display formats. This makes it possible for viewers to simultaneously grasp the emotional state of the person who filmed the content while viewing it.
[0532] A "user terminal" refers to a device that has the function of acquiring and analyzing digital content data.
[0533] "Digital content data" is a general term that refers to information acquired electronically, such as video and audio.
[0534] "Emotional information" refers to data that analyzes a user's emotional state and expresses it as numerical values or categories.
[0535] An "information processing device" refers to a device that stores, converts, and generates distribution links for received digital content data and emotional information.
[0536] "Display format" refers to the specific form in which digital content data is converted into a format that can be played on the viewer's device.
[0537] "Communication link" refers to an access method created to provide digital content data and emotional information to viewers via the internet.
[0538] "Display device user" refers to an individual or organization that receives and views digital content data using the provided link.
[0539] "Internal program" refers to software embedded in a display device that reads communication links and plays digital content data.
[0540] The system of this invention consists of a user terminal, a server, and a display device. The user uses a user terminal, such as a smartphone or camera device, to acquire event and daily digital content data. The user terminal is equipped with software called an emotion engine, which analyzes and recognizes emotional information from the user's facial expressions and voice.
[0541] The user's terminal compresses digital content data using compression technologies such as the H.264 codec and generates emotional information. This compressed data and emotional information are transmitted to a server via the internet. The server stores the received digital content data in its storage and converts it into various display formats such as MP4 and WebM. This makes it easy to view the content on different devices.
[0542] The server generates a communication link based on the converted digital content data. This link, along with sentiment information, is sent to the user of the display device. Common notification methods include email and social media. Upon receiving the link, the user opens it via a smart TV or streaming device, and the content is played by an internal program, while the sentiment information is displayed as an overlay.
[0543] As a concrete example, consider a scenario where a user films a family picnic and wants to share the joyful moment with relatives. The user sends the video along with identified emotion information, specifically "joy," to a server. The server generates a communication link based on the picnic video data and emotion information and notifies the relatives. The relatives open this link on their television and share the joy of that moment while watching the picnic video.
[0544] An example of a prompt for a generative AI model would be: "Describe in detail how a user can share a moment they experienced with others through digital content data and visually convey emotional information."
[0545] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0546] Step 1:
[0547] Users acquire digital content data using smartphones or camera devices. The input consists of video and audio from events and daily life, which are analyzed by an emotion engine. The device analyzes the user's facial expressions and tone of voice to recognize emotional information. The output of this step is raw digital content data and its associated emotional information.
[0548] Step 2:
[0549] The device compresses the acquired digital content data using the H.264 codec. The input is the video and audio data obtained in step 1. This data is compressed to reduce its size, and the compressed digital content data is generated as output. In addition, emotional information is converted to JSON format.
[0550] Step 3:
[0551] The terminal sends compressed digital content data and sentiment information in JSON format to the server via the internet. The input for this step is the compressed data and sentiment information, which are sent to the server for use in subsequent processing. The output is the data stored in a format usable by the server.
[0552] Step 4:
[0553] The server stores the received digital content data and sentiment information in storage. The input consists of compressed data and sentiment information sent from the terminal. The server converts the digital content data into multiple display formats (e.g., MP4, WebM). The output is digital content data that can be viewed on different devices.
[0554] Step 5:
[0555] The server generates a communication link from the converted digital content data and notifies the user of this link and sentiment information. The input is the converted digital content data and sentiment information. The server generates a link based on this and notifies the user via email, social media, etc. The output is the notification of the communication link and sentiment information.
[0556] Step 6:
[0557] The user of the display device receives the notified link and opens it using an internal program on their smart TV or streaming device. This plays the digital content data and displays emotional information visually. The input is the communication link, and the output is the viewable digital content data and the display of emotional information.
[0558] (Application Example 2)
[0559] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0560] In today's content sharing environment, there is a lack of means to provide a richer experience that not only allows viewers to watch videos but also simultaneously conveys the emotions of the person who filmed them. Furthermore, there is a need for a simpler and faster way to achieve this.
[0561] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0562] In this invention, the server includes means for transmitting video image data and emotion data acquired from a user terminal via the Internet, means for converting the received video image data into multiple broadcast formats, and means for generating a streaming identifier based on the converted video image data and emotion data and notifying the recipient. This allows viewers to perceive the emotions of the person filming the video simultaneously with the video itself, enabling them to enjoy a more interactive experience.
[0563] A "user terminal" is an electronic device that an individual can carry and use to record and transmit video and audio.
[0564] "Motion image data" refers to data that records visual image information in digital format.
[0565] "Compression" is a process that reduces data size to enable more efficient transmission or storage.
[0566] "Emotional data" refers to digital information that indicates a user's psychological state, obtained by analyzing their facial expressions and tone of voice.
[0567] The "Internet" is a communication framework that connects networks around the world, enabling the transmission and reception of information.
[0568] A "server device" is a computing device connected to a network for receiving, processing, and transmitting large amounts of data.
[0569] A "broadcast format" is a standard digital format for efficiently distributing moving images and videos.
[0570] A "streaming identifier" is information such as a link or hash that identifies specific video data online and facilitates its distribution.
[0571] A "recipient" is an end-user or device that receives data from a sender.
[0572] A "video output device" is a display device that reproduces received video and audio data visually and audibly.
[0573] To implement this invention, a system is constructed that includes a user terminal, a server device, and a video output device for the receiver.
[0574] Program processing
[0575] 1. User terminal operation
[0576] The user terminal is a portable device such as a smartphone or smart glasses, used to acquire video and image data and user emotion data. This terminal is equipped with a camera and microphone, and simultaneously records video and audio, while also having emotion analysis software (e.g., TensorFlow, Amazon Rekognition) installed. This allows the user's facial expressions and voice tone to be analyzed along with the video, generating emotion data.
[0577] 2. Role of the Server Device
[0578] The server device receives video and emotion data transmitted from user terminals and processes them into an efficient format. Specifically, it converts the data into multiple broadcast formats (e.g., H.264, MP4) and generates a streaming identifier. This allows recipients to easily access the data through the identifier.
[0579] 3. Receiver's video output device
[0580] The recipient's video output device is a smart TV or streaming player, which streams video data and sentiment data using an identifier transmitted from the server. The device has an internal program that analyzes the received identifier and displays the video and sentiment information in real time.
[0581] Specific example
[0582] For example, suppose a user wants to record their performance at a festival and share their joy and surprise with friends. In this case, the user's device records the user's emotions along with the video and uploads the data to a server. The server converts the data to a predetermined format and generates an identifier for sending to friends. By using that identifier to play the video on a smart TV, the friend can share the festival footage along with the user's joy and surprise.
[0583] Example of a prompt
[0584] "A student is recording a science experiment show at a science museum. Capture his smile and surprised expression, and suggest the best way for him to share his excitement with his friends."
[0585] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0586] Step 1:
[0587] The user uses a device to record video and audio data. The device uses its camera and microphone to acquire video and audio as input data. From this input data, emotion analysis software installed on the device analyzes the user's facial expressions and voice tone to generate emotion data. The emotion data is output as a numerical value or category indicating the user's psychological state (e.g., joy, surprise).
[0588] Step 2:
[0589] The terminal compresses the generated video and emotion data. A codec (e.g., H.264) is used for compression to reduce data size and facilitate transmission. At this stage, the input is uncompressed video and emotion data, and the output is a compressed data file.
[0590] Step 3:
[0591] The terminal transmits compressed video data and emotion data to the server over the internet. The input is the data compressed in step 2, and the output is the reception of the data on the server. This transmission is performed using a secure and fast protocol (e.g., HTTPS).
[0592] Step 4:
[0593] The server converts the received video data into multiple broadcast formats. The input data is the compressed data received in step 3, and the output is the data converted to a different format (e.g., MP4, WebM). A media conversion library (e.g., FFmpeg) is used for this process.
[0594] Step 5:
[0595] The server generates a streaming identifier based on the converted video and sentiment data and notifies the recipient. The input is the converted data and sentiment data, and the output is provided as a streaming link or code accessible to the recipient. A unique hash or URL is used to generate the identifier.
[0596] Step 6:
[0597] The recipient receives the notified identifier and starts streaming playback on the video output device. The device takes the streaming link as input and displays the video data and emotion data in real time. The output is the video and emotion information that the recipient can see. Through this process, content including emotions is delivered from the user to the recipient.
[0598] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0599] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0600] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0601] [Fourth Embodiment]
[0602] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0603] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0604] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0605] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0606] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0607] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0608] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0609] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0610] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0611] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0612] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0613] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0614] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0615] The system of the present invention consists of a user device, a server device, and a viewer's video display device. The user device functions as a smartphone or camera device and acquires video data of events or travel. The user uses a dedicated application on the terminal to select the captured video data, compress it, and transmit it to the server device via the internet.
[0616] The server device stores the received video data in storage and converts it to a television format. Based on this converted data, it generates a streaming link and notifies the viewer. The viewer uses this link to watch in real time on their home video display device.
[0617] The viewer's video display device functions as a smart TV or streaming device, accessing the notified link through an internal application. This application reads the link, receives streaming data from the server, and displays the video on the TV screen.
[0618] For example, if a user films their grandchild at a school sports day and wants to share the video with grandparents living far away, the user uploads the video using a dedicated app. The server converts the data and sends a generated link to the grandparents' email address. The grandparents can then open the link in a TV application and enjoy watching their grandchild's video on a large screen.
[0619] In this way, the system of the present invention can be easily operated even by viewers unfamiliar with technology, and makes it possible to share what family members are doing on a large screen even if they are far away.
[0620] The following describes the processing flow.
[0621] Step 1:
[0622] Users use their smartphones or camera devices to film events or travels. They then launch a dedicated application and select the filmed video data.
[0623] Step 2:
[0624] The terminal receives the selected video data and compresses it to efficiently transmit it while maintaining data quality. The compressed video data is then sent to the server via the internet.
[0625] Step 3:
[0626] The server stores video data received via the internet in its storage. The server analyzes the received data and performs a process to convert it into multiple television formats (e.g., H.264, MP4).
[0627] Step 4:
[0628] The server generates a streaming link accessible to viewers based on the converted video data. The server then notifies viewers of this link via email or messaging services.
[0629] Step 5:
[0630] Viewers activate their home video display device and check the provided link. They open the internal application on their smart TV or streaming device.
[0631] Step 6:
[0632] The device reads links entered or scanned by the viewer using its internal application. The device receives streaming data from the server and plays the video on the display device's screen in real time.
[0633] (Example 1)
[0634] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0635] Conventional video distribution systems have presented challenges in that they require complex operations and specialized knowledge to efficiently deliver video information captured by users to recipients in remote locations. Therefore, there is a need for a simple and effective video distribution system that is easy to use even for users unfamiliar with technology or elderly recipients.
[0636] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0637] In this invention, the server includes means for compressing video information acquired from a terminal and transmitting it to a data processing device via an information network; means for the data processing device to convert the received video information into multiple visual output formats; and means for generating connection information for distribution based on the converted video information and notifying the recipient. This makes it possible for users to easily operate and distribute videos in real time to family and friends in remote locations without technical constraints.
[0638] A "terminal" refers to an information processing device used by users to acquire video information during events or trips, compress it, and send it to a data processing device.
[0639] "Video information" refers to visual recording data acquired by a device, which is compressed and then distributed.
[0640] An "information network" refers to a means of communication for transmitting video information from a terminal to a data processing device, and includes a wide range of networks such as the internet.
[0641] A "data processing device" refers to a device that receives video information transmitted from a terminal, converts it into multiple visual output formats, and generates connection information for distribution.
[0642] A "visual output format" is a data format used when video information is distributed and visualized, enabling proper display on different devices.
[0643] "Distribution connection information" refers to the links or metadata necessary to deliver the converted video information to recipients, and is generated to make it accessible to recipients.
[0644] A "receiver" refers to a person or device that receives content by visualizing video information transmitted from a device.
[0645] This invention relates to a system for easily distributing video information to recipients in remote locations. Users acquire video information using smartphones or digital cameras as terminals. These terminals are equipped with dedicated applications for efficiently compressing the acquired video information. Standard algorithms such as H.264 and HEVC are used for compression, and after compression, the data is transmitted to a data processing device via an information network.
[0646] The data processing unit functions as a server, saving received video information to a storage system. After saving, software such as FFmpeg is used to convert the video information into multiple visual output formats. This ensures that the information can be properly visualized on any recipient's display device. The converted video information is then used to generate connection information for distribution, which is then notified to the recipient via email or messaging services.
[0647] The recipient uses the verified connection information to visualize the video information in real time using a display device equipped with the internal program, such as a smart TV or streaming device. This allows users to share videos with family and friends in different locations and enjoy them on a larger screen.
[0648] For example, if a user films their child's school play and wants to share it with relatives living abroad, they would upload the video to a server via a dedicated app. An example of a prompt message would be, "Please tell me the steps to take to share a video of my child's concert with relatives living abroad." This creates an environment where even recipients unfamiliar with technology can easily visualize the process.
[0649] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0650] Step 1:
[0651] Users acquire video information using their smartphones or digital cameras. The acquired videos are previewed within a dedicated application, and the user selects which videos to send. The recorded video files are used as input, and the selected videos are prepared within the application as output.
[0652] Step 2:
[0653] The device compresses the selected video. Algorithms such as H.264 and HEVC are used for compression. The input is the selected video file, and the output is a compressed file. This compression reduces the video file size and improves transfer efficiency.
[0654] Step 3:
[0655] The user's terminal sends a compressed video file to a data processing device via an information network. The input is the compressed video file, and the output is the completion of the file transfer to the server. Security is ensured by using protocols such as HTTPS during this process.
[0656] Step 4:
[0657] The server saves the received video to its storage. The input is the transmitted video file, and the output is the saved file that exists in storage. Data integrity is verified during this saving process.
[0658] Step 5:
[0659] The server converts the saved video into a visual output format. Software such as FFmpeg is used to convert it to television formats and other streaming formats. The input is the saved video file, and the output consists of multiple converted visual output format files.
[0660] Step 6:
[0661] The server generates connection information for distribution based on the converted video and notifies the recipient. The input includes the converted video data and the recipient's contact information, while the output is the generated streaming link information and confirmation of its transmission. This information is sent to the recipient via email.
[0662] Step 7:
[0663] The recipient's display device streams the video from its internal program using the notified connection information. The input is the transmitted streaming link information, and based on this, it receives real-time streaming data from the server. The output is the video playback on the recipient's display device. The recipient can easily view the video.
[0664] (Application Example 1)
[0665] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0666] Conventional information sharing systems struggled to quickly and efficiently visualize high-quality information data on multiple display devices. Furthermore, image quality degradation and delays during data compression and conversion processes were problematic, highlighting the need for improved user experience. In particular, there was a need to create a user-friendly system for elderly individuals and those unfamiliar with technology.
[0667] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0668] In this invention, the server includes means for efficiently compressing information data acquired from a user terminal and transmitting it to a central device via a data communication network; means for the central device to convert the received information data into multiple display formats; and means for providing dynamically generated access information based on the converted information data and notifying the information requester. This enables even information requesters unfamiliar with the technology to quickly and efficiently visualize high-quality information data.
[0669] A "user terminal" is a device that has the function of acquiring information data, compressing it efficiently, and transmitting it.
[0670] "High-efficiency compression" refers to a process that reduces data size while maintaining data quality, thereby improving communication speed.
[0671] A "data communication network" is a network used to send and receive information data between devices located in remote locations.
[0672] A "central device" is a device that has the function of converting received information data into multiple display formats and dynamically generating access information.
[0673] "Display format" refers to various formats and protocols used to visualize information data.
[0674] "Dynamically generated access information" refers to information that is generated on the spot by the information requester to provide the necessary links and authentication information to visualize the information data.
[0675] An "information requester" is a person or device that visualizes information data using the provided access information.
[0676] A "display device" is a device that has the function of receiving and visualizing information data.
[0677] An "internal processing program" is software that reads access information and performs visualization of the information data.
[0678] The system of the present invention consists of a user terminal, a central unit, and a display device. The user terminal is a device equipped with a camera and communication functions, and is responsible for acquiring information data and compressing it efficiently. Specifically, the user takes videos of daily life or special events with the user terminal, and the size of the video data is reduced using H.264 or H.265 video compression technology. The compressed information data is then transmitted to the central unit via a data communication network.
[0679] The central system is built on cloud platforms such as Amazon Web Services (AWS) or Microsoft Azure. The central system converts the received information data into various display formats (e.g., MP4 or HLS). This conversion process ensures the information data is formatted to be properly visualized on diverse display devices. The central system then creates dynamically generated access information to notify the information requester, including links and authentication credentials for accessing the information data.
[0680] The information requester uses this access information to visualize the information data on their display device. The display device is a smart TV or tablet, which reads the access information provided by an internal processing program and receives and displays the streaming data.
[0681] As a concrete example, consider a scenario where a user records a family trip and wants to share this video with family members who live far away. The user compresses the video on their device and sends it to a central device. The central device performs the necessary conversions and generates a link that allows for smooth visualization on display devices. This link is sent to family members via email or message, allowing them to enjoy the trip in real time on their own displays.
[0682] Examples of prompts to input into a generative AI model:
[0683] "Please explain in detail how to design a system that efficiently compresses video footage taken during family trips and streams it to family members in real time."
[0684] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0685] Step 1:
[0686] The user acquires information data using their device. Specifically, the user captures video with a camera device and acquires that video data within the device. The input is the captured video data, and the output is the raw video data.
[0687] Step 2:
[0688] The terminal efficiently compresses information data. The user terminal compresses video data using H.264 or H.265 codecs to reduce data size. The input is the acquired video data, and the output is the compressed video data. This compression process improves communication speed and enables efficient data transmission.
[0689] Step 3:
[0690] The terminal transmits compressed information data to the central device via the data communication network. The user terminal transmits the compressed video data to the server via the internet. The input is the compressed video data, and the output is the status of completion of data transmission to the server.
[0691] Step 4:
[0692] The server converts the received data into multiple display formats. The central unit converts the received video data into MP4 or HLS format. The input is the compressed data sent to the server, and the output is the converted video data. This ensures compatibility with various display devices.
[0693] Step 5:
[0694] The server dynamically generates access information based on the converted information data and notifies the information requester. The server generates a streaming link based on the converted data and notifies the specified information requester of the link information via email or message. The input is the converted video data, and the output is the generated streaming link.
[0695] Step 6:
[0696] The information requester's display device visualizes the information data using the notified link information. The information requester reads the received link using the display device's internal processing program and streams the video data. The input is the received streaming link, and the output is the video playback on the display device.
[0697] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0698] The system of the present invention consists of a user device, a server device, a viewer's video display device, and an emotion engine. The user device functions as a smartphone or camera device and not only acquires video data of events and trips, but also has an emotion engine that analyzes the user's facial expressions and voice tone to recognize emotions. The user device selects the captured video data, compresses the video and emotion data, and transmits it to the server device via the internet.
[0699] The server device stores the received video data and emotion data in storage and converts the video to multiple television formats (e.g., H.264, MP4). Based on the converted video data, the server generates a streaming link and notifies the viewer along with the emotion data.
[0700] The viewer's video display device functions as a smart TV or streaming device, displaying notified links and sentiment data in real time using an internal application. This allows viewers to understand the emotional state of the person filming while simultaneously watching the video.
[0701] For example, if a user films their grandchild's sports day and wants to share their emotions (for example, joy) with their grandparents, the user uploads the video along with emotion data recognized by an emotion engine to a server via a dedicated application. The server generates a streaming link combining the sports day video and information indicating the joyful emotion, and sends it to the grandparents. The grandparents can open this link on their TV and watch their grandchild's video while also sharing the emotions of the moment.
[0702] This invention allows for not only viewing of images, but also the simultaneous transmission of the photographer's emotions, enabling a deeper experience and communication.
[0703] The following describes the processing flow.
[0704] Step 1:
[0705] The user uses a smartphone or camera device to record video during events or trips. Simultaneously, an emotion engine built into the user's device analyzes the user's facial expressions and tone of voice to recognize their emotions.
[0706] Step 2:
[0707] The device acquires the captured video data and emotion data recognized by the emotion engine. The device compresses the video data and prepares it to be sent to the server with the emotion data added.
[0708] Step 3:
[0709] The device sends compressed video and emotional data to the server via the internet. Once the transmission is complete, it sends instructions to the server to proceed to the next process.
[0710] Step 4:
[0711] The server stores the received video and emotion data in storage. The server analyzes this video data and converts it into a format suitable for television.
[0712] Step 5:
[0713] The server generates a streaming link based on the converted video data. Simultaneously, it stores the sentiment data sent by the user, associating it with the streaming link.
[0714] Step 6:
[0715] The server notifies viewers of the generated streaming link and associated sentiment data. This notification is sent via email or messaging services.
[0716] Step 7:
[0717] Viewers activate their home video display device and check the notified link. Viewers use an internal application on their smart TV or streaming device to enter or select the streaming link.
[0718] Step 8:
[0719] The device receives input from the viewer and receives streaming data from the server. The device plays the video in real time, displaying emotion data that matches the video data on the display screen.
[0720] (Example 2)
[0721] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0722] Traditional digital content distribution systems only deliver digital content such as video and audio, making it difficult to convey the emotions and intentions of the creator to the viewer. There is a growing need for viewers to understand the emotions behind the content and gain a deeper experience.
[0723] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0724] In this invention, the server includes means for analyzing digital content data acquired from a user terminal and recognizing emotional information, means for transmitting the digital content data and emotional information to an information processing device via the Internet, and means for converting the digital content data received by the information processing device into multiple display formats. This makes it possible for viewers to simultaneously grasp the emotional state of the person who filmed the content while viewing it.
[0725] A "user terminal" refers to a device that has the function of acquiring and analyzing digital content data.
[0726] "Digital content data" is a general term that refers to information acquired electronically, such as video and audio.
[0727] "Emotional information" refers to data that analyzes a user's emotional state and expresses it as numerical values or categories.
[0728] An "information processing device" refers to a device that stores, converts, and generates distribution links for received digital content data and emotional information.
[0729] "Display format" refers to the specific form in which digital content data is converted into a format that can be played on the viewer's device.
[0730] "Communication link" refers to an access method created to provide digital content data and emotional information to viewers via the internet.
[0731] "Display device user" refers to an individual or organization that receives and views digital content data using the provided link.
[0732] "Internal program" refers to software embedded in a display device that reads communication links and plays digital content data.
[0733] The system of this invention consists of a user terminal, a server, and a display device. The user uses a user terminal, such as a smartphone or camera device, to acquire event and daily digital content data. The user terminal is equipped with software called an emotion engine, which analyzes and recognizes emotional information from the user's facial expressions and voice.
[0734] The user's terminal compresses digital content data using compression technologies such as the H.264 codec and generates emotional information. This compressed data and emotional information are transmitted to a server via the internet. The server stores the received digital content data in its storage and converts it into various display formats such as MP4 and WebM. This makes it easy to view the content on different devices.
[0735] The server generates a communication link based on the converted digital content data. This link, along with sentiment information, is sent to the user of the display device. Common notification methods include email and social media. Upon receiving the link, the user opens it via a smart TV or streaming device, and the content is played by an internal program, while the sentiment information is displayed as an overlay.
[0736] As a concrete example, consider a scenario where a user films a family picnic and wants to share the joyful moment with relatives. The user sends the video along with identified emotion information, specifically "joy," to a server. The server generates a communication link based on the picnic video data and emotion information and notifies the relatives. The relatives open this link on their television and share the joy of that moment while watching the picnic video.
[0737] An example of a prompt for a generative AI model would be: "Describe in detail how a user can share a moment they experienced with others through digital content data and visually convey emotional information."
[0738] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0739] Step 1:
[0740] Users acquire digital content data using smartphones or camera devices. The input consists of video and audio from events and daily life, which are analyzed by an emotion engine. The device analyzes the user's facial expressions and tone of voice to recognize emotional information. The output of this step is raw digital content data and its associated emotional information.
[0741] Step 2:
[0742] The device compresses the acquired digital content data using the H.264 codec. The input is the video and audio data obtained in step 1. This data is compressed to reduce its size, and the compressed digital content data is generated as output. In addition, emotional information is converted to JSON format.
[0743] Step 3:
[0744] The terminal sends compressed digital content data and sentiment information in JSON format to the server via the internet. The input for this step is the compressed data and sentiment information, which are sent to the server for use in subsequent processing. The output is the data stored in a format usable by the server.
[0745] Step 4:
[0746] The server stores the received digital content data and sentiment information in storage. The input consists of compressed data and sentiment information sent from the terminal. The server converts the digital content data into multiple display formats (e.g., MP4, WebM). The output is digital content data that can be viewed on different devices.
[0747] Step 5:
[0748] The server generates a communication link from the converted digital content data and notifies the user of this link and sentiment information. The input is the converted digital content data and sentiment information. The server generates a link based on this and notifies the user via email, social media, etc. The output is the notification of the communication link and sentiment information.
[0749] Step 6:
[0750] The user of the display device receives the notified link and opens it using an internal program on their smart TV or streaming device. This plays the digital content data and displays emotional information visually. The input is the communication link, and the output is the viewable digital content data and the display of emotional information.
[0751] (Application Example 2)
[0752] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0753] In today's content sharing environment, there is a lack of means to provide a richer experience that not only allows viewers to watch videos but also simultaneously conveys the emotions of the person who filmed them. Furthermore, there is a need for a simpler and faster way to achieve this.
[0754] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0755] In this invention, the server includes means for transmitting video image data and emotion data acquired from a user terminal via the Internet, means for converting the received video image data into multiple broadcast formats, and means for generating a streaming identifier based on the converted video image data and emotion data and notifying the recipient. This allows viewers to perceive the emotions of the person filming the video simultaneously with the video itself, enabling them to enjoy a more interactive experience.
[0756] A "user terminal" is an electronic device that an individual can carry and use to record and transmit video and audio.
[0757] "Motion image data" refers to data that records visual image information in digital format.
[0758] "Compression" is a process that reduces data size to enable more efficient transmission or storage.
[0759] "Emotional data" refers to digital information that indicates a user's psychological state, obtained by analyzing their facial expressions and tone of voice.
[0760] The "Internet" is a communication framework that connects networks around the world, enabling the transmission and reception of information.
[0761] A "server device" is a computing device connected to a network for receiving, processing, and transmitting large amounts of data.
[0762] A "broadcast format" is a standard digital format for efficiently distributing moving images and videos.
[0763] A "streaming identifier" is information such as a link or hash that identifies specific video data online and facilitates its distribution.
[0764] A "recipient" is an end-user or device that receives data from a sender.
[0765] A "video output device" is a display device that reproduces received video and audio data visually and audibly.
[0766] To implement this invention, a system is constructed that includes a user terminal, a server device, and a video output device for the receiver.
[0767] Program processing
[0768] 1. User terminal operation
[0769] The user terminal is a portable device such as a smartphone or smart glasses, used to acquire video and image data and user emotion data. This terminal is equipped with a camera and microphone, and simultaneously records video and audio, while also having emotion analysis software (e.g., TensorFlow, Amazon Rekognition) installed. This allows the user's facial expressions and voice tone to be analyzed along with the video, generating emotion data.
[0770] 2. Role of the Server Device
[0771] The server device receives video and emotion data transmitted from user terminals and processes them into an efficient format. Specifically, it converts the data into multiple broadcast formats (e.g., H.264, MP4) and generates a streaming identifier. This allows recipients to easily access the data through the identifier.
[0772] 3. Receiver's video output device
[0773] The recipient's video output device is a smart TV or streaming player, which streams video data and sentiment data using an identifier transmitted from the server. The device has an internal program that analyzes the received identifier and displays the video and sentiment information in real time.
[0774] Specific example
[0775] For example, suppose a user wants to record their performance at a festival and share their joy and surprise with friends. In this case, the user's device records the user's emotions along with the video and uploads the data to a server. The server converts the data to a predetermined format and generates an identifier for sending to friends. By using that identifier to play the video on a smart TV, the friend can share the festival footage along with the user's joy and surprise.
[0776] Example of a prompt
[0777] "A student is recording a science experiment show at a science museum. Capture his smile and surprised expression, and suggest the best way for him to share his excitement with his friends."
[0778] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0779] Step 1:
[0780] The user uses a device to record video and audio data. The device uses its camera and microphone to acquire video and audio as input data. From this input data, emotion analysis software installed on the device analyzes the user's facial expressions and voice tone to generate emotion data. The emotion data is output as a numerical value or category indicating the user's psychological state (e.g., joy, surprise).
[0781] Step 2:
[0782] The terminal compresses the generated video and emotion data. A codec (e.g., H.264) is used for compression to reduce data size and facilitate transmission. At this stage, the input is uncompressed video and emotion data, and the output is a compressed data file.
[0783] Step 3:
[0784] The terminal transmits compressed video data and emotion data to the server over the internet. The input is the data compressed in step 2, and the output is the reception of the data on the server. This transmission is performed using a secure and fast protocol (e.g., HTTPS).
[0785] Step 4:
[0786] The server converts the received video data into multiple broadcast formats. The input data is the compressed data received in step 3, and the output is the data converted to a different format (e.g., MP4, WebM). A media conversion library (e.g., FFmpeg) is used for this process.
[0787] Step 5:
[0788] The server generates a streaming identifier based on the converted video and sentiment data and notifies the recipient. The input is the converted data and sentiment data, and the output is provided as a streaming link or code accessible to the recipient. A unique hash or URL is used to generate the identifier.
[0789] Step 6:
[0790] The recipient receives the notified identifier and starts streaming playback on the video output device. The device takes the streaming link as input and displays the video data and emotion data in real time. The output is the video and emotion information that the recipient can see. Through this process, content including emotions is delivered from the user to the recipient.
[0791] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0792] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0793] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0794] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0795] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0796] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0797] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0798] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0799] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0800] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0801] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0802] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0803] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0804] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0805] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0806] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0807] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0808] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0809] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0810] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0811] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0812] The following is further disclosed regarding the embodiments described above.
[0813] (Claim 1)
[0814] The video data acquired from the user's device is compressed,
[0815] Means for transmitting the aforementioned video data to a server device via the Internet,
[0816] The server device includes means for converting video data received into multiple television formats,
[0817] A means for generating a streaming link based on the converted video data and notifying the viewer,
[0818] A means for streaming and displaying video data to a viewer's video display device via the aforementioned notified link,
[0819] A system that includes this.
[0820] (Claim 2)
[0821] The system according to claim 1, which includes a setting that the user device automatically compresses video data.
[0822] (Claim 3)
[0823] The system according to claim 1, wherein the viewer's video display device has means for reading the link using an internal application.
[0824] "Example 1"
[0825] (Claim 1)
[0826] Compress the video information obtained from the device.
[0827] Means for transmitting the aforementioned video information to a data processing device via an information network,
[0828] The data processing device includes means for converting video information received by the data processing device into multiple visual output formats,
[0829] A means for generating connection information for distribution based on the converted video information and notifying the recipient,
[0830] A means for distributing and visualizing video information to a recipient's display device via the aforementioned notified connection information,
[0831] A system that includes this.
[0832] (Claim 2)
[0833] The system according to claim 1, which includes a setting that the terminal automatically compresses video information.
[0834] (Claim 3)
[0835] The system according to claim 1, wherein the receiver's display device has means for reading the connection information using an internal program.
[0836] "Application Example 1"
[0837] (Claim 1)
[0838] The information data acquired from the user terminal is compressed with high efficiency.
[0839] Means for transmitting the aforementioned information data to a central device via a data communication network,
[0840] The central device includes means for converting the information data received into multiple display formats,
[0841] A means for providing dynamically generated access information based on the converted information data and notifying the information requester,
[0842] A means for supplying and visualizing information data to the information requester's display device through the provided access information,
[0843] A system that includes this.
[0844] (Claim 2)
[0845] The system according to claim 1, which includes a setting that the user terminal automatically compresses information data with high efficiency.
[0846] (Claim 3)
[0847] The system according to claim 1, wherein the information requester's display device is equipped with means for reading the access information using an internal processing program.
[0848] "Example 2 of combining an emotion engine"
[0849] (Claim 1)
[0850] A means of analyzing digital content data acquired from user terminals and recognizing emotional information,
[0851] Means for transmitting the aforementioned digital content data and emotional information to an information processing device via the Internet,
[0852] The information processing device includes means for converting digital content data received by the information processing device into multiple display formats,
[0853] A means for generating a communication link based on the converted digital content data and notifying the user of the display device along with emotional information,
[0854] A means for distributing and displaying digital content data on a display device of a display device user via the aforementioned notified link,
[0855] A system that includes this.
[0856] (Claim 2)
[0857] The system according to claim 1, which includes a setting that the user terminal automatically analyzes digital content data and recognizes emotional information.
[0858] (Claim 3)
[0859] The system according to claim 1, wherein the display device of the user of the display device has means for reading the link using an internal program.
[0860] "Application example 2 when combining with an emotional engine"
[0861] (Claim 1)
[0862] Compress the video data acquired from the user terminal.
[0863] Means for transmitting the aforementioned video data and emotion data to a server device via the Internet,
[0864] The server device includes means for converting video data received into multiple broadcast formats,
[0865] A means for generating a streaming identifier based on the converted video data and emotion data and notifying the recipient,
[0866] A means for streaming video data and emotion data to a recipient's video output device via the notified identifier and displaying them,
[0867] A system that includes this.
[0868] (Claim 2)
[0869] The system according to claim 1, which includes a setting that the user terminal automatically compresses video data and performs sentiment analysis.
[0870] (Claim 3)
[0871] The system according to claim 1, wherein the receiver's video output device has means for reading the identifier using an internal program. [Explanation of symbols]
[0872] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. The video data acquired from the user's device is compressed, Means for transmitting the aforementioned video data to a server device via the Internet, The server device includes means for converting video data received into multiple television formats, A means for generating a streaming link based on the converted video data and notifying the viewer, A means for streaming and displaying video data to a viewer's video display device via the aforementioned notified link, A system that includes this.
2. The system according to claim 1, which includes a setting that the user device automatically compresses video data.
3. The system according to claim 1, wherein the viewer's video display device has means for reading the link using an internal application.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A