System

By employing smart glasses and image recognition algorithms to overlay relevant information in augmented reality format, the video viewing experience is enriched with real-time trivia and details, addressing the limitations of traditional viewing experiences.

JP2025071043APending Publication Date: 2025-05-02SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024182288
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-19
Filing Date
2024-10-17
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

Traditional video viewing experiences lack the ability for viewers to easily obtain information or trivia related to video scenes, resulting in a shallow understanding and lack of engagement.

Method used

A system using user-wearable display devices like smart glasses, which applies image recognition algorithms to identify buildings and clothing in video scenes, retrieves related information from databases or AI models, converts it into augmented reality format, and overlays it on the user's field of vision.

Benefits of technology

This solution enhances the video viewing experience by providing viewers with detailed information and trivia in real-time, deepening their understanding and engagement with the content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025071043000001_ABST
    Figure 2025071043000001_ABST
Patent Text Reader

Abstract

To provide a system that makes a moving image viewing experience rich and full of information through a display unit worn by a user of smart glasses.SOLUTION: A system is to provide, in an augmented reality format, information related to moving images viewed by using a display unit worn by a user, and includes: means that applies an image recognition algorithm for performing image recognition on the moving images viewed by using the display unit, and identifies at least a structure or the person's clothes in a scene in the moving images; means that, on the basis of the result of the image recognition, acquires relevant information on the structure or the person's clothes from a database or a generative AI model; means that converts the format of the acquired relevant information into the augmented reality format; means that transmits information on the augmented reality format obtained through the conversion to the display unit; and means that causes the display unit to display the information on the augmented reality format superimposed on the user's field of view.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including a description and related instruction sentence regarding the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] JP 2022-180282 A Summary of the Invention [Problem to be solved by the invention]

[0004] Traditional video viewing experiences have the challenge that viewers cannot easily obtain information or trivia related to scenes in the video, making it difficult for them to gain a deeper understanding or find interesting information. [Means for solving the problem]

[0005] The present invention provides a means for providing a rich and informative video viewing experience through a user-worn display device such as smart glasses. Specifically, the present invention includes the following systems:

[0006] A system for providing information in an augmented reality format related to a video viewed by a user using a display device worn by the user, the system comprising: means for applying an image recognition algorithm to the video viewed on the display device to identify at least a building or a person's clothing in a scene of the video; means for obtaining related information related to the building or the person's clothing from a database or a generative AI model based on the results of the image recognition; means for converting the obtained related information into an augmented reality format; means for transmitting the converted augmented reality format information to the display device; and means for displaying the augmented reality format information by the display device, superimposing it on the user's field of vision.

[0007] This allows viewers to enjoy augmented reality-style information while watching videos, providing real-time details and trivia related to scenes in the video, enriching the viewing experience for viewers and providing deeper understanding and interesting information.

[0008] "Image recognition" is a technology that identifies specific objects or features from video frames or images and retrieves related information.

[0009] "Related information" refers to information or trivia that may be of interest to viewers, such as detailed information about buildings or clothes worn by characters in video scenes, historical background, or brand or styling information.

[0010] "Augmented reality" refers to technology and methods that overlay virtual information and objects onto the real-world environment. By overlaying information on the viewer's field of vision through smart glasses, it provides relevant information and trivia in real time. [Brief description of the drawings]

[0011] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Diagram 2]1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. FIG. [Diagram 3] FIG. 11 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Diagram 5] FIG. 13 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 13 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 13 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] 4 is a sequence diagram showing a process flow of the data processing system according to the first embodiment. FIG. [Figure 12] 11 is a sequence diagram showing a process flow of the data processing system in application example 1. FIG. [Figure 13] FIG. 11 is a sequence diagram showing the flow of processing of the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 11 is a sequence diagram showing the flow of processing in the data processing system in application example 2 when combined with an emotion engine. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0012] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0013] First, the terms used in the following description will be explained.

[0014] In the following embodiments, a signed processor (hereinafter simply referred to as a "processor") may be one arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be one type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), etc.

[0015] In the following embodiments, a signed RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by the processor.

[0016] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0017] In the following embodiments, a communication I / F (Interface) with a code is an interface including a communication processor and an antenna. The communication I / F controls communication between multiple computers. An example of a communication standard applied to the communication I / F is a wireless communication standard including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0018] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. In addition, in this specification, the same idea as "A and / or B" is also applied when three or more things are expressed by connecting them with "and / or."

[0019] [First embodiment]

[0020] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0021] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0022] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a wide area network (WAN) and / or a local area network (LAN).

[0023] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0024] The reception device 38 includes a touch panel 38A and a microphone 38B, and receives user input. The touch panel 38A detects contact with a pointer (e.g., a pen or a finger) to receive user input by the touch of the pointer. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0025] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (e.g., voice and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs voice according to instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, an aperture, and a shutter, and an imaging element such as a Complementary Metal-Oxide-Semiconductor (CMOS) image sensor or a Charge Coupled Device (CCD) image sensor.

[0026] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for transmitting and receiving various types of information between the processor 46 and the processor 28 via the network 54.

[0027] FIG. 2 shows an example of main functions of the data processing device 12 and the smart device 14.

[0028] As shown in Fig. 2, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32. The specific process program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific process program 56 from the storage 32, and executes the read specific process program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific process program 56 executed on the RAM 30.

[0029] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0030] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores a reception output program 60. The reception output program 60 is used together with the specific processing program 56 by the data processing system 10. The processor 46 reads out the reception output program 60 from the storage 50, and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0031] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0032] An embodiment for implementing the present invention is a system including the following elements.

[0033] 1. Server: Acquires video frames and applies image recognition algorithms to recognize buildings and clothes worn by characters. Based on the recognition results, retrieves related information from a database and processes it in AR format. The database may be one that has been prepared in advance, or may be one that has been searched for and temporarily stored during processing using the Internet, or may be one that has been generated using a data generation model 58.

[0034] 2. Smart glasses (terminal): Receives AR information sent from the server and displays it over the user's field of view. Viewers can view the AR information through the smart glasses.

[0035] In a specific example, a user wears smart glasses and watches a movie. The server retrieves frames of each scene of the movie and uses image recognition algorithms to recognize buildings and clothes worn by characters. For example, if the building appearing in the movie scene is the Eiffel Tower, the server retrieves detailed information and historical background about the Eiffel Tower from the database. Similarly, it retrieves information about the clothes worn by the characters.

[0036] The server processes the acquired related information in AR format and sends it to the smart glasses. The smart glasses display the received AR information, allowing users to enjoy the AR information while watching the movie. For example, when the Eiffel Tower appears in a scene in the movie, the smart glasses display detailed information about the Eiffel Tower and its historical background in AR. The same is true for the clothes worn by the characters.

[0037] The above is an embodiment of the present invention, and by viewing AR information provided through the smart glasses, a user can have a rich and informative video viewing experience.

[0038] The process flow of each embodiment will be described below.

[0039] Step 1: The server acquires frames of a video, which allows the server to understand the scene of the video.

[0040] Step 2: The server processes the captured frames with image recognition algorithms, specifically analyzing features and patterns in the images to recognize buildings and the clothes worn by characters.

[0041] Step 3: The server retrieves related information from the database based on the recognition results, such as detailed information about the building, historical background, and the brand and styling of the clothing worn by the character.

[0042] Step 4: The server processes the acquired related information in AR format and sends it to the user's smart glasses. Specifically, the server uses a communication protocol to package the related information as text and images as AR data and send it to the user's smart glasses.

[0043] Step 5: The smart glasses (terminal) displays the received AR information. Specifically, by overlaying related information on the smart glasses display, the user can view the AR information while watching the video.

[0044] Example 1

[0045] Next, a description will be given of Example 1. In the following description, the data processing device 12 is referred to as a "server" and the smart device 14 is referred to as a "terminal."

[0046] In conventional video viewing systems, it was difficult to obtain detailed information related to the video the user was watching in real time and provide it visually. In addition, there was a lack of a method to instantly provide detailed information about buildings and people's clothing while watching the video, and there were limited ways to enrich the video viewing experience.

[0047] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0048] In this invention, the server includes a means for continuously acquiring frames of a video being viewed by a display device worn by a user in order to provide information related to the video being viewed by the display device in an augmented reality format, a means for applying an image recognition algorithm to the acquired frames of the video to identify buildings and clothing of people, a means for acquiring related information related to the buildings and clothing of people from a database or a generative AI model based on the result of the image recognition, a means for converting the acquired related information into an augmented reality format, a means for transmitting the converted augmented reality information to the display device, and a means for displaying the augmented reality information by the display device in a way that is superimposed on the user's field of view. This allows the user to display detailed information acquired in real time superimposed on the video being viewed, thereby significantly improving the video viewing experience.

[0049] A "user" is a person who uses the system to watch videos and receive the augmented reality information provided.

[0050] A "display device" is a device worn by a user to display information superimposed on the user's field of vision, and includes smart glasses and the like.

[0051] "Video" refers to continuous video data such as television or movies that are viewed by users.

[0052] A "frame" is an individual still image that makes up a video.

[0053] An "image recognition algorithm" is a type of computer program that identifies objects based on captured frames.

[0054] "Buildings" refers to buildings and structures that appear in the video.

[0055] "People" refers to people who appear in the video.

[0056] "Clothing" refers to the clothes and accessories worn by a person.

[0057] "Related information" refers to detailed information or background information related to the clothing of a structure or person identified by the image recognition algorithm.

[0058] A "database" is a data storage system that stores related information and from which a server can retrieve the information.

[0059] A "generative AI model" is an artificial intelligence program that generates the necessary information by inputting a prompt sentence.

[0060] The term "augmented reality" refers to a format in which additional information is overlaid on a real-world image.

[0061] "Server" means a computer system that captures video frames, applies image recognition algorithms, captures related information, and processes and transmits the augmented reality information.

[0062] A "prompt sentence" is an input sentence that requests the generative AI model to generate specific information.

[0063] An embodiment of the present invention is a system that provides, in an augmented reality format, information related to a video being viewed by a user using a display device worn by the user.

[0064] Hardware and Software

[0065] The hardware includes the following:

[0066] High-performance GPU server (e.g. general-purpose GPU server)

[0067] Display devices (e.g. smart glasses)

[0068] The software includes the following:

[0069] OpenCV library for acquiring and processing video frames

[0070] TENSORFLOW® Framework for Image Recognition Algorithms

[0071] Unity® and ARKit® / ARCore® for generating augmented reality information

[0072] Database management systems for information retrieval and generative AI models (e.g., general generative AI models)

[0073] Specific explanation of operation

[0074] The server continuously captures frames of the video being viewed on the display device. For example, the server opens a movie file and uses OpenCV to extract frames every second. The server then applies a TensorFlow model to the captured frames to perform image recognition and identify buildings and people's clothing.

[0075] Relevant information about the identified structures and clothing is retrieved using a database or a generative AI model. For example, the server inputs a prompt such as "What is the history of the Eiffel Tower?" into the generative AI model and retrieves the generated information.

[0076] The acquired related information is converted into an augmented reality format using Unity or ARKit. Specifically, the generated text information is processed as a 3D object and converted into a format suitable for a display device. The converted augmented reality information is transmitted to the display device via Wi-Fi (registered trademark) or Bluetooth (registered trademark).

[0077] The terminal (display device) displays the received augmented reality information by overlaying it on the user's field of view. For example, the display device displays information about the Eiffel Tower in the user's field of view, and when the user moves his or her head, the information moves accordingly.

[0078] Examples

[0079] Consider a scenario where a user is watching a movie through a display device worn by the user. The server captures frames of each scene of the movie every second and runs an image recognition algorithm on the captured frames. Based on the recognition results, the server retrieves detailed information about the Eiffel Tower from a database or generates information by inputting a prompt sentence such as "Tell me the history of the Eiffel Tower" into a generative AI model.

[0080] Using the prompt sentence "What is the history of the Eiffel Tower?", the AI ​​model retrieves detailed information. The information is converted into an augmented reality format and sent to the display device. The display device overlays the retrieved information on the user's field of view, allowing the user to experience the movie scene and related details in an augmented reality format.

[0081] The above system enables users to receive detailed information obtained in real time in combination with the video they are watching, greatly improving the video viewing experience.

[0082] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0083] Step 1: Grab a video frame

[0084] The server retrieves each frame sequentially from the video source: it opens the video file and uses OpenCV to extract a frame every second.

[0085] Input: Video file

[0086] Output: Still images (frames)

[0087] What happens: The server reads the stream of the video file and captures frames at regular intervals (e.g. every second).

[0088] Step 2: Applying image recognition algorithms

[0089] The server applies image recognition algorithms to the captured frames, using TensorFlow to identify buildings and people's clothing.

[0090] Input: Still image (frame)

[0091] Output: Recognition results (location and type of buildings and clothing)

[0092] Specific operation: The server runs a deep learning model (e.g., YOLO or ResNet) on the captured frames to identify objects.

[0093] Step 3: Get relevant information

[0094] Based on the recognition results, the server retrieves relevant information from a database or generative AI model, for example, detailed information about a building or a person's clothing.

[0095] Input: Recognition result

[0096] Output: Related information

[0097] Specific operation: Based on the recognition results, the server sends a query to a database to obtain the required information, or inputs a prompt sentence such as "Please tell me the history of the Eiffel Tower" into a generative AI model and obtains the generated information.

[0098] Step 4: Convert to AR format

[0099] The server converts the relevant information into an augmented reality format, using Unity or ARKit to process the information into a format suitable for the user.

[0100] Input: Related information

[0101] Output: Augmented reality information

[0102] Specific operation: The server generates the text information and images it acquires as 3D objects and animations, and processes them into an augmented reality format.

[0103] Step 5: Submit AR information

[0104] The server then sends the processed augmented reality information to the display device. Data is transferred using Wi-Fi or Bluetooth.

[0105] Input: Augmented reality information

[0106] Output: The data sent.

[0107] Specific operation: The server specifies the IP address of the display device and transmits the augmented reality information.

[0108] Step 6: View AR information

[0109] The terminal (display device) displays the received augmented reality information by overlaying it on the user's field of vision using a device such as smart glasses.

[0110] Input: Data sent

[0111] Output: Augmented reality display in user's field of view

[0112] What it does: The device caches the data it receives and displays it in real time according to the user's field of view. Specifically, it overlays information about the Eiffel Tower into the user's field of view, adjusting the information in response to the user's gaze and head movements.

[0113] (Application example 1)

[0114] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0115] In conventional viewing experiences, it is difficult for users to immediately obtain information that interests them, and it is also difficult to physically experience the product, making it difficult to stimulate purchasing motivation. Furthermore, the viewing experience is often interrupted because it takes time to search and collect information. This leads to a poor user experience and a decrease in satisfaction.

[0116] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0117] In this invention, the server includes a means for image recognition of an object, a means for acquiring related information via a database or the Internet, and a means for processing the acquired related information in an AR format. This allows the user to obtain related information in real time without interrupting the viewing experience, and the product information and usage examples displayed on the smart glasses are expected to enrich the overall experience and increase purchasing motivation.

[0118] A "user" is someone who uses the smart glasses to have a viewing experience.

[0119] "Smart glasses" are glasses-type devices that have the ability to overlay augmented reality (AR) information on visual information.

[0120] A "viewing experience" is the act of a user wearing smart glasses and observing information or images.

[0121] "Real-time" refers to near-instant processing and results.

[0122] "Related information" is data that includes detailed information or use cases related to objects recognized during a viewing experience.

[0123] "AR" is an abbreviation for Augmented Reality, a technology that overlays digital information on the real world.

[0124] The term "system" refers to a collection of a series of devices and software in which a plurality of means in the present invention function in cooperation with one another.

[0125] An "object" is an object that is captured by the smart glasses' camera during the viewing experience and is the subject of image recognition.

[0126] "Image recognition" is the technology of identifying specific objects from visual data.

[0127] A "database" is a collection of information for storing and managing related information.

[0128] "Internet" is a general term for networks used to search and retrieve information.

[0129] "AR format" refers to a form of information display that uses augmented reality technology.

[0130] This invention relates to a system that provides information related to a viewing experience in real time using AR when a user is viewing an image using smart glasses. The main components and operations for specifically implementing the invention will be described below.

[0131] First, the smart glasses (terminal) are used by the user during the viewing experience and are equipped with a camera and a display. The smart glasses have the function of capturing visual information and transmitting it to a server.

[0132] The server performs image recognition of the object from the received video data. Specific image recognition algorithms such as TensorFlow and YOLO are used for this image recognition. After recognizing the object, the server retrieves related information. Related information can be retrieved from databases such as MongoDB, or the required information can be searched via the Internet.

[0133] The acquired information is converted to AR format on the server and sent to the smart glasses. The AR information is generated using AR libraries such as Vuforia (registered trademark) and Wikitude. The smart glasses display the received AR information overlaid on the user's field of view. This allows the user to obtain relevant information in real time without interrupting the viewing experience.

[0134] The key point of the present invention is that the smart glasses can display detailed information and usage examples of products directly in the user's field of vision, enriching the overall viewing experience and increasing the user's purchasing motivation, for example, allowing a boutique customer to check the details of a garment without trying it on, or a furniture store customer to instantly learn the details of a piece of furniture.

[0135] The specific hardware and software used is as follows:

[0136] Smart Glasses: Examples of smart glasses include Google® Glass® and Microsoft® HoloLens®.

[0137] Server: A cloud-based server with high-performance processing power, such as Amazon Web Services (AWS®) or Microsoft Azure®.

[0138] Image recognition algorithms: TensorFlow and YOLO.

[0139] Database: MongoDB(registered trademark) or MySQL(registered trademark).

[0140] AR libraries: Vuforia and Wikitude.

[0141] An example of a prompt might be:

[0142] "Create an application that uses AR to display detailed information about a particular product in a virtual store."

[0143] The above is a specific embodiment of the present invention, which allows users to enrich their viewing experience and obtain relevant information in real time.

[0144] The flow of the specific process in the application example 1 will be described with reference to FIG.

[0145] Step 1:

[0146] The smart glasses (terminal) capture video data during the user's viewing experience. The camera in the smart glasses acquires the video data in real time, encodes the data in JPEG format, and sends it to the server.

[0147] Input: Video data of the user's viewing experience

[0148] Output: Encoded data in JPEG format (sent to server)

[0149] How it works: The camera in the smart glasses captures video and converts the frames into JPEG format, which is then sent to the server.

[0150] Step 2:

[0151] The server performs image recognition of the object from the received video data. It uses image recognition algorithms such as TensorFlow and YOLO to identify the object from the transmitted video data.

[0152] Input: Encoded data in JPEG format

[0153] Output: Information about the recognized object

[0154] How it works: The server receives the video data and applies image recognition algorithms to identify objects, such as clothing or furniture.

[0155] Step 3:

[0156] The server retrieves related information based on the results of image recognition, searching for and retrieving the necessary information via databases such as MongoDB or the Internet.

[0157] Input: Information about the recognized object

[0158] Output: Related information (details, use cases, historical background)

[0159] Specific operation: The server collects detailed information about the recognized object from a database or the Internet. For example, in the case of clothing, information about the material and brand is obtained.

[0160] Step 4:

[0161] The server processes the relevant information obtained into AR format, using AR libraries such as Vuforia and Wikitude to convert it into a visually easy-to-understand format.

[0162] Input: Related information (details, use cases, historical background)

[0163] Output: AR format data

[0164] Specific operation: The server generates data for AR display based on the related information collected. For example, it processes the acquired clothing information so that it can be displayed as a 3D model.

[0165] Step 5:

[0166] The server sends the processed AR information to the smart glasses, which then overlay the information onto the user's field of vision.

[0167] Input: AR format data

[0168] Output: AR information displayed on smart glasses

[0169] Specific operation: The server transmits the processed AR information to the smart glasses, which then display the information in real time, allowing users to obtain AR information that is visually intuitive.

[0170] Furthermore, an emotion engine that estimates the emotion of the user may be combined. That is, the identification processing unit 290 may estimate the emotion of the user using the emotion identification model 59, and perform identification processing using the emotion of the user.

[0171] An embodiment for implementing the present invention is a system including the following elements.

[0172] 1. Server: Obtains video frames and applies image recognition algorithms to recognize buildings and clothes worn by characters. Based on the recognition results, retrieves related information from a database and processes it in AR format. Furthermore, combines an emotion engine to recognize the user's emotions and adjusts related information and AR display. The database may be one that has been prepared in advance, or may be one that has been searched for and temporarily stored during processing using the Internet, or may be one that has been generated using a data generation model 58.

[0173] 2. Smart glasses (terminal): Receives AR information sent from the server and displays it over the user's field of view. The smart glasses recognize the user's emotions and adjust the relevant information and AR display accordingly.

[0174] As a specific example, a user wears smart glasses and watches a movie. The server retrieves frames of each scene in the movie and uses an image recognition algorithm to recognize buildings and clothes worn by characters. At the same time, it uses an emotion engine to recognize the user's emotions. For example, if the user is smiling, the emotion engine analyzes the emotion and adjusts it to provide relevant information and AR displays tailored to the user's interests and preferences.

[0175] The server processes the acquired related information in AR format and sends it to the smart glasses. The smart glasses display the received AR information, so that the user can enjoy the AR information while watching the movie. At the same time, the smart glasses recognize the user's emotions and adjust the related information and AR display according to the user's emotions. For example, if the user is smiling, the smart glasses adjust the related information and AR display to be more enjoyable.

[0176] The above is an embodiment of the present invention, and by viewing AR information provided through the smart glasses, a user can have a rich, informative and emotionally tailored video viewing experience.

[0177] The process flow of each embodiment will be described below.

[0178] Step 1: The server acquires frames of a video, which allows the server to understand the scene of the video.

[0179] Step 2: The server processes the captured frames with image recognition algorithms, specifically analyzing features and patterns in the images to recognize buildings and the clothes worn by characters.

[0180] Step 3: The server retrieves related information from the database based on the recognition results, such as detailed information about the building, historical background, and the brand and styling of the clothing worn by the character.

[0181] Step 4: The server uses the emotion engine to recognize the user's emotions. Specifically, it analyzes emotions from the user's facial expressions and voice, and adjusts them to provide relevant information and AR displays that match the user's interests and preferences.

[0182] Step 5: The server combines the acquired related information with the emotion recognition results and processes the AR information. Specifically, the related information is packaged as AR data in the form of text and images, and the AR display is adjusted according to the user's emotions.

[0183] Step 6: The server sends the processed AR information to the user's smart glasses. Specifically, the server sends the AR data to the smart glasses using a communication protocol.

[0184] Step 7: The smart glasses (terminal) displays the received AR information. Specifically, by overlaying related information and AR display on the smart glasses display, the user can view the AR information while watching the video.

[0185] Step 8: The smart glasses recognize the user's emotions. Specifically, they analyze the user's emotions from their facial expressions and voice, and adjust the relevant information and AR display according to the user's emotions.

[0186] Example 2

[0187] Next, a description will be given of Example 2. In the following description, the data processing device 12 is referred to as a "server" and the smart device 14 is referred to as a "terminal."

[0188] In conventional video viewing systems using visual devices, the means to provide information related to the video in real time are limited, limiting the user's viewing experience. In addition, there is a lack of technology to adjust related information taking into account the user's emotions. As a result, the information obtained while watching a video may not match the user's emotional state or interests.

[0189] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0190] In this invention, the server includes a means for performing image recognition of at least a structure or clothing worn by a character in a scene of a video, a means for acquiring related information from a data storage device based on a result of the image recognition, and a means for recognizing a user's emotion using an emotion engine and adjusting the related information and the augmented reality display in order to process the acquired related information in an augmented reality format. This makes it possible to provide related information in real time in augmented reality in accordance with the user's emotional state while watching a video.

[0191] A "visual device" is a digital device that enables a user to view images and videos, including, for example, smart glasses and head-mounted displays.

[0192] "Video" refers to visual content such as television, movies, video streaming services, etc., including content that a user views through a visual device.

[0193] "Related information" refers to knowledge or data related to a particular element in a video (e.g. a building, or the clothing worn by a character), including details, historical context, styling information, etc.

[0194] "Augmented reality" refers to a technology that overlays digital information on the user's field of vision, providing real-time information to the user's field of vision.

[0195] "Image recognition" refers to the use of computer vision techniques to identify and classify objects and features within a video scene.

[0196] "Data storage device" refers to a system that stores relevant information and makes it available for search and retrieval as needed, including databases and cloud storage.

[0197] An "emotion engine" refers to algorithms and technologies that analyze human facial expressions and behavior to infer their emotional state.

[0198] "Augmented reality processing" refers to the process of adjusting and transforming acquired data to fit the user's field of view, including adding visual effects and optimizing the way it is displayed.

[0199] "Real-time" refers to the fact that processing and display occur instantly at the moment the user begins viewing.

[0200] An embodiment of this system will now be described, which mainly comprises three parties: a server, a visual device (such as smart glasses), and a user.

[0201] Hardware and Software Configuration

[0202] server

[0203] The server acts as the main processing unit and uses the following hardware and software:

[0204] Hardware: A general server machine equipped with a high-performance CPU, GPU, and large memory capacity

[0205] software:

[0206] FFmpeg: A library for capturing video frames

[0207] TensorFlow: A library for implementing image recognition algorithms

[0208] Data storage device: Database management system such as MySQL (registered trademark) or MongoDB

[0209] Emotion engine: Emotion recognition services such as Microsoft's Emotion API

[0210] AR development tools: Unity, Vuforia

[0211] Data processing and calculation

[0212] The server first captures video data frame by frame using FFmpeg. Next, it performs image recognition on these frames using TensorFlow. Based on the recognized information on buildings and clothes, it retrieves related information from databases such as MySQL and MongoDB. After that, it analyzes the user's emotions using Emotion API and processes the related information in an augmented reality format using Unity or Vuforia.

[0213] Smart glasses (terminal)

[0214] The smart glasses have the function of receiving augmented reality information sent from a server and displaying it over the user's field of vision. This is achieved by using the following hardware and software:

[0215] Hardware: Smart glasses with built-in camera, display sensor and processor

[0216] Software: Custom application for displaying augmented reality information

[0217] User Actions

[0218] The user wears the smart glasses and watches the video. Facial expression information captured by the smart glasses while watching the video is sent to the server in real time. The server analyzes this using an emotion engine, creates augmented reality information appropriate to the user's emotional state, and sends it back to the smart glasses. While watching the video, the user can visually obtain detailed information about buildings and clothes in real time.

[0219] Examples

[0220] For example, if a user is watching a historical movie, the server will recognize a building from the movie scene and retrieve the building's historical information from the database. At the same time, if the user's facial expression indicates that they are moved, the server will send augmented reality information with an emotional effect to the smart glasses. While watching the movie, the user can enjoy information such as the background and story of the building, making the experience richer.

[0221] Examples of prompt statements

[0222] Prompt: A user is watching a movie that contains a scene depicting a certain historical building. The user's emotion is recognized as being touched. Please suggest the most suitable AR information and its presentation effect that should be displayed in this situation.

[0223] The above is an embodiment of the present invention.

[0224] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0225] The flow of this system's program processing

[0226] Step 1: Grab a video frame

[0227] The server captures frames from the video the user is watching, using the FFmpeg library.

[0228] Input: The video the user is watching

[0229] Output: Captured frame image

[0230] Specific behavior:

[0231] The server runs the FFmpeg command to capture a frame every second: ffmpeg -i input.mp4 -vf "fps=1" frame_%04d.png.

[0232] Step 2: Applying image recognition algorithms

[0233] The server uses TensorFlow to perform image recognition on the captured frames, which allows it to recognize buildings and the clothes worn by characters.

[0234] Input: The captured frame image

[0235] Output: Recognized buildings and clothing information

[0236] Specific behavior:

[0237] The server loads a TensorFlow trained model and runs each frame through the model to obtain recognition results.

[0238] Example: recognition_result = tensorflow_model.predict(frame)

[0239] Step 3: Retrieving relevant information from the database

[0240] Based on the results of image recognition, the server retrieves relevant information about buildings and clothes from a database using MySQL and MongoDB.

[0241] Input: Recognized building and clothing information

[0242] Output: Related information (details, historical background, styling information)

[0243] Specific behavior:

[0244] The server retrieves the information by executing an SQL query like SELECT FROM building_info WHERE name = 'recognized_building'.

[0245] Step 4: Emotion recognition by emotion engine

[0246] The server uses an emotion engine such as Microsoft's Emotion API to recognize the user's emotions and receives the user's facial expression data from the smart glasses.

[0247] Input: User facial expression data

[0248] Output: The perceived emotional state of the user.

[0249] Specific behavior:

[0250] The server sends the user's facial expression data to the Emotion API and obtains the emotion analysis results.

[0251] Example: emotion_result = emotion_api.analyze(user_face_image)

[0252] Step 5: Processing the AR information

[0253] The server uses Unity or Vuforia to generate and process AR information based on the acquired related information and the user's emotional state.

[0254] Input: relevant information, the user's emotional state

[0255] Output: Processed AR information

[0256] Specific behavior:

[0257] The server generates 3D objects and effects in Unity and creates the AR display.

[0258] For users who are impressed, add specific effects (eg, fireworks or music).

[0259] Step 6: Sending information to the smart glasses

[0260] The server sends the processed AR information to the smart glasses, updating the information in real time using protocols such as WebSocket.

[0261] Input: Processed AR information

[0262] Output: Sending information to smart glasses

[0263] Specific behavior:

[0264] The server sends information using websocket.send(ar_information).

[0265] Step 7: Display and emotional feedback on smart glasses

[0266] The device (smart glasses) receives the AR information from the server and displays it over the user's field of view. At the same time, it recognizes the user's emotions again and updates the information as necessary.

[0267] Input: AR information sent from the server

[0268] Output: AR information displayed in the user's field of view, and the user's emotion data obtained again.

[0269] Specific behavior:

[0270] The smart glasses display the AR information received through a display application in the field of view.

[0271] The built-in camera is used to capture the user's facial expressions in real time and the data is sent to the server.

[0272] (Application example 2)

[0273] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0274] In conventional video viewing, there was a lack of a way to provide users with detailed information about the content in real time. In addition, there was no mechanism to adjust related information according to the user's emotions, so the viewing experience was uniform and it was not possible to provide information optimized for each individual user.

[0275] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means. In this invention, the server includes a means for performing image recognition of at least a building or clothes worn by a character in a scene of a video, a means for acquiring related information from a database based on the result of the image recognition, a means for processing the acquired related information in an AR format and transmitting it to the user's smart device, a means for recognizing the user's face using a camera of the smart device and analyzing emotions, and a means for adjusting the display content of the related information based on the analyzed emotions. This makes it possible to provide the user with an optimized viewing experience by providing the user with related information in real time and further adjusting the information according to the user's emotions.

[0276] A "smart device" is a device such as a smartphone, smart glasses, or a head-mounted display that is equipped with a camera to display images and recognize the user's face.

[0277] "Video" refers to dynamic video content, such as movies, television programs, and online streaming content.

[0278] "Building" refers to any structure or structure that appears in the video.

[0279] "Clothes worn by characters" refers to the clothing worn by people or characters appearing in the video.

[0280] An "image recognition means" is an algorithm or software that extracts information from frames of video and identifies specific objects or people.

[0281] "Related information" is supplemental data provided to users, such as detailed information about buildings or clothing, historical background, design information, etc.

[0282] A "database" is a collection of related information that can be stored and retrieved as needed.

[0283] "Means for processing in AR format" refers to a technology that uses augmented reality (AR) technology to superimpose related information on the user's field of vision.

[0284] A "transmitting means" is a communication means or protocol for sending data from the server to the smart device.

[0285] "Means for recognizing faces and analyzing emotions" refers to technology that captures the user's face with a camera and determines their emotional state through facial expression analysis, etc.

[0286] "Means for adjusting display content" refers to technology for changing the content or style of information displayed in AR based on recognized emotions.

[0287] The present invention is a system that provides related information in real time based on the content of a video when a user watches the video using a smart device, and further adjusts the information according to the user's emotions. The following describes how to implement the system in detail.

[0288] First, the server captures frames from the video and uses image recognition technology to identify buildings and character clothing. Image recognition is achieved using OpenCV, a commonly used image processing library, and an AI framework equipped with a deep learning model. The server also retrieves relevant information from a database and processes it into an AR format. The database pre-stores relevant information, but it is also possible to temporarily retrieve information via the Internet during processing.

[0289] Next, the smart device (e.g., smartphone, smart glasses) receives the AR information sent from the server and displays it over the user's field of view. This uses augmented reality (AR) technology such as Google ARCore. In addition, the camera of the smart device is used to recognize the user's face and perform emotion analysis. A pre-trained emotion recognition model (e.g., a deep learning model built using Keras) is used for emotion analysis. Once the emotion is analyzed, the display content of the AR information is adjusted based on the results. For example, if the user smiles, the color and display style of the related information are adjusted to be brighter.

[0290] As a concrete example, while a user is watching their favorite movie on their smartphone, behind-the-scenes and historical information about buildings and costumes that appear in the movie are displayed in real time in AR, and the information is adjusted according to the user's emotions: if the user smiles, the information display becomes more cheerful and bright.

[0291] The specific program processing for realizing this system is as follows.

[0292] 1. The server captures frames of video and uses image recognition technology to identify buildings and character clothing.

[0293] 2. Retrieve relevant information from the database, process it into AR format and send it to the smart device.

[0294] 3. Recognize the user's face using the smart device's camera and perform emotion analysis. The emotion recognition model used is built using Keras.

[0295] 4. The display content of the AR information is adjusted based on the emotion recognition results. For example, the color and style are adjusted to be brighter when the person is smiling.

[0296] 5. The smart device will then display this tailored information, providing the user with an improved viewing experience.

[0297] An example of a prompt for a generative AI model is as follows:

[0298] Suppose a user is watching a video on their smartphone. Information about buildings and costumes that appear in the movie scenes is retrieved in real time and displayed in AR format. Furthermore, the smartphone camera recognizes the user's emotions, and if the user smiles, the display of related information is adjusted to be more enjoyable.

[0299] The flow of the specific process in the application example 2 will be described with reference to FIG.

[0300] Step 1:

[0301] The server captures the video stream from the user's smart device in real-time. It extracts frames of the video and uses them for analysis. The input is the video stream and the output is the individual frames. This is done using an image processing library like OpenCV.

[0302] Step 2:

[0303] The server applies image recognition techniques to the extracted frames. Specifically, it uses a deep learning model to identify buildings and character clothing in the frames. The input is the captured frame, and the output is the identified building and clothing information. This is done using deep learning frameworks such as TensorFlow and PyTorch.

[0304] Step 3:

[0305] The server retrieves information related to the identified buildings or clothes from a database. The input is the identified object and the output is the related information. This can be done using a SQL or NoSQL database.

[0306] Step 4:

[0307] The server processes the retrieved relevant information into an augmented reality (AR) format. The input is the relevant information, and the output is the processed information for AR. This is achieved using an AR library such as Google ARCore.

[0308] Step 5:

[0309] The server sends the processed AR information to the user's smart device. The input is the processed AR information, and the output is the data sent to the user's smart device. This is done using REST APIs and WebSockets.

[0310] Step 6:

[0311] The device uses a camera to recognize the user's face. It then uses a facial recognition algorithm to analyze the user's emotions. The input is the camera image, and the output is the user's emotional information. Keras and OpenCV are used for emotion recognition.

[0312] Step 7:

[0313] The terminal adjusts the AR information received from the server based on the user's emotion information. For example, if the user is smiling, the color and style of the displayed information is adjusted to be brighter. The input is the user's emotion information and the processed AR information, and the output is the adjusted AR information.

[0314] Step 8:

[0315] The device overlays the adjusted AR information onto the user's field of view. The input is the adjusted AR information, and the output is the AR content displayed in the user's field of view. This is done using an AR framework such as ARCore or ARKit.

[0316] As a concrete example, when a user is watching a movie, detailed information about the buildings and clothes that appear in the scene is displayed, and if the user smiles, the information is adjusted to be more pleasant.

[0317] Suppose a user is watching a video on their smartphone. Information about buildings and costumes that appear in the movie scenes is retrieved in real time and displayed in AR format. Furthermore, the smartphone camera recognizes the user's emotions, and if the user smiles, the display of related information is adjusted to be more enjoyable.

[0318] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires a voice indicating a user input for the result of the specific processing. The control unit 46A transmits the voice data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[0319] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by making a neural network perform deep learning. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating a voice, text data indicating a text, and image data indicating an image is input. The data generation model 58 performs inference on the input inference data according to the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0320] In the above embodiment, an example was given in which the specific process was performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0321] [Second embodiment]

[0322] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0323] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0324] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a wide area network (WAN) and / or a local area network (LAN).

[0325] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0326] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs the voice according to instructions from the processor 46.

[0327] Camera 42 is a small digital camera equipped with an optical system including a lens, an aperture, and a shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (e.g., an imaging range defined by an angle of view equivalent to the width of the field of vision of an average healthy person).

[0328] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for transmitting and receiving various types of information between the processor 46 and the processor 28 via the network 54. The transmission and reception of various types of information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is performed in a secure state.

[0329] Fig. 4 shows an example of main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0330] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32, and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0331] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0332] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50, and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0333] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal".

[0334] An embodiment for implementing the present invention is a system including the following elements.

[0335] 1. Server: Acquires video frames and applies image recognition algorithms to recognize buildings and clothes worn by characters. Based on the recognition results, retrieves related information from a database and processes it in AR format. The database may be one that has been prepared in advance, or may be one that has been searched for and temporarily stored during processing using the Internet, or may be one that has been generated using a data generation model 58.

[0336] 2. Smart glasses (terminal): Receives AR information sent from the server and displays it over the user's field of view. Viewers can view the AR information through the smart glasses.

[0337] In a specific example, a user wears smart glasses and watches a movie. The server retrieves frames of each scene of the movie and uses image recognition algorithms to recognize buildings and clothes worn by characters. For example, if the building appearing in the movie scene is the Eiffel Tower, the server retrieves detailed information and historical background about the Eiffel Tower from the database. Similarly, it retrieves information about the clothes worn by the characters.

[0338] The server processes the acquired related information in AR format and sends it to the smart glasses. The smart glasses display the received AR information, allowing users to enjoy the AR information while watching the movie. For example, when the Eiffel Tower appears in a scene in the movie, the smart glasses will display detailed information about the Eiffel Tower and its historical background in AR, as well as the clothes worn by the characters.

[0339] The above is an embodiment of the present invention, and by viewing AR information provided through the smart glasses, a user can have a rich and informative video viewing experience.

[0340] The process flow will be explained below.

[0341] Step 1: The server acquires frames of a video, which allows the server to understand the scene of the video.

[0342] Step 2: The server processes the captured frames with image recognition algorithms, specifically analyzing features and patterns in the images to recognize buildings and the clothes worn by characters.

[0343] Step 3: The server retrieves related information from the database based on the recognition results, such as detailed information about the building, historical background, and the brand and styling of the clothing worn by the character.

[0344] Step 4: The server processes the acquired related information in AR format and sends it to the user's smart glasses. Specifically, the server uses a communication protocol to package the related information as text and images as AR data and send it to the user's smart glasses.

[0345] Step 5: The smart glasses (terminal) displays the received AR information. Specifically, by overlaying related information on the smart glasses display, the user can view the AR information while watching the video.

[0346] Example 1

[0347] Next, a description will be given of Example 1. In the following description, the data processing device 12 is referred to as a "server" and the smart glasses 214 are referred to as a "terminal".

[0348] In conventional video viewing systems, it was difficult to obtain detailed information related to the video the user was watching in real time and provide it visually. In addition, there was a lack of a method to instantly provide detailed information about buildings and people's clothing while watching the video, and there were limited ways to enrich the video viewing experience.

[0349] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0350] In this invention, the server includes a means for continuously acquiring frames of a video being viewed by a display device worn by a user in order to provide information related to the video being viewed by the display device in an augmented reality format, a means for applying an image recognition algorithm to the acquired frames of the video to identify buildings and clothing of people, a means for acquiring related information related to the buildings and clothing of people from a database or a generative AI model based on the result of the image recognition, a means for converting the acquired related information into an augmented reality format, a means for transmitting the converted augmented reality information to the display device, and a means for displaying the augmented reality information by the display device in a way that is superimposed on the user's field of view. This allows the user to display detailed information acquired in real time superimposed on the video being viewed, thereby significantly improving the video viewing experience.

[0351] A "user" is a person who uses the system to watch videos and receive the augmented reality information provided.

[0352] A "display device" is a device worn by a user to display information superimposed on the user's field of vision, and includes smart glasses and the like.

[0353] "Video" refers to continuous video data such as television or movies that are viewed by users.

[0354] A "frame" is an individual still image that makes up a video.

[0355] An "image recognition algorithm" is a type of computer program that identifies objects based on captured frames.

[0356] "Buildings" refers to buildings and structures that appear in the video.

[0357] "People" refers to people who appear in the video.

[0358] "Clothing" refers to the clothes and accessories worn by a person.

[0359] "Related information" refers to detailed information or background information related to the clothing of a structure or person identified by the image recognition algorithm.

[0360] A "database" is a data storage system that stores related information and from which a server can retrieve the information.

[0361] A "generative AI model" is an artificial intelligence program that generates the necessary information by inputting a prompt sentence.

[0362] The term "augmented reality" refers to a format in which additional information is overlaid on a real-world image.

[0363] "Server" means a computer system that captures video frames, applies image recognition algorithms, captures related information, and processes and transmits the augmented reality information.

[0364] A "prompt sentence" is an input sentence that requests the generative AI model to generate specific information.

[0365] An embodiment of the present invention is a system that provides, in an augmented reality format, information related to a video being viewed by a user using a display device worn by the user.

[0366] Hardware and Software

[0367] The hardware includes the following:

[0368] High-performance GPU server (e.g. general-purpose GPU server)

[0369] Display devices (e.g. smart glasses)

[0370] The software includes the following:

[0371] OpenCV library for acquiring and processing video frames

[0372] TensorFlow Framework for Image Recognition Algorithms

[0373] Unity and ARKit / ARCore for generating augmented reality information

[0374] Database management systems for information retrieval and generative AI models (e.g., general generative AI models)

[0375] Specific explanation of operation

[0376] The server continuously captures frames of the video being viewed on the display device. For example, the server opens a movie file and uses OpenCV to extract frames every second. The server then applies a TensorFlow model to the captured frames to perform image recognition and identify buildings and people's clothing.

[0377] Relevant information about the identified structures and clothing is retrieved using a database or a generative AI model. For example, the server inputs a prompt such as "What is the history of the Eiffel Tower?" into the generative AI model and retrieves the generated information.

[0378] The acquired related information is converted into an augmented reality format using Unity or ARKit. Specifically, the generated text information is processed as a 3D object and converted into a format suitable for the display device. The converted augmented reality information is sent to the display device via Wi-Fi or Bluetooth.

[0379] The terminal (display device) displays the received augmented reality information by overlaying it on the user's field of view. For example, the display device displays information about the Eiffel Tower in the user's field of view, and when the user moves his or her head, the information moves accordingly.

[0380] Examples

[0381] Consider a scenario where a user is watching a movie through a display device worn by the user. The server captures frames of each scene of the movie every second and runs an image recognition algorithm on the captured frames. Based on the recognition results, the server retrieves detailed information about the Eiffel Tower from a database or generates information by inputting a prompt sentence such as "Tell me the history of the Eiffel Tower" into a generative AI model.

[0382] Using the prompt sentence "What is the history of the Eiffel Tower?", the AI ​​model retrieves detailed information. The information is converted into an augmented reality format and sent to the display device. The display device overlays the retrieved information on the user's field of view, allowing the user to experience the movie scene and related details in an augmented reality format.

[0383] The above system enables users to receive detailed information obtained in real time in combination with the video they are watching, greatly improving the video viewing experience.

[0384] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0385] Step 1: Grab a video frame

[0386] The server retrieves each frame sequentially from the video source: it opens the video file and uses OpenCV to extract a frame every second.

[0387] Input: Video file

[0388] Output: Still images (frames)

[0389] What happens: The server reads the stream of the video file and captures frames at regular intervals (e.g. every second).

[0390] Step 2: Applying image recognition algorithms

[0391] The server applies image recognition algorithms to the captured frames, using TensorFlow to identify buildings and people's clothing.

[0392] Input: Still image (frame)

[0393] Output: Recognition results (location and type of buildings and clothing)

[0394] Specific operation: The server runs a deep learning model (e.g., YOLO or ResNet) on the captured frames to identify objects.

[0395] Step 3: Get relevant information

[0396] Based on the recognition results, the server retrieves relevant information from a database or generative AI model, for example, detailed information about a building or a person's clothing.

[0397] Input: Recognition result

[0398] Output: Related information

[0399] Specific operation: Based on the recognition results, the server sends a query to a database to obtain the required information, or inputs a prompt sentence such as "Please tell me the history of the Eiffel Tower" into a generative AI model and obtains the generated information.

[0400] Step 4: Convert to AR format

[0401] The server converts the relevant information into an augmented reality format, using Unity or ARKit to process the information into a format suitable for the user.

[0402] Input: Related information

[0403] Output: Augmented reality information

[0404] Specific operation: The server generates the text information and images it acquires as 3D objects and animations, and processes them into an augmented reality format.

[0405] Step 5: Submit AR information

[0406] The server then sends the processed augmented reality information to the display device. Data is transferred using Wi-Fi or Bluetooth.

[0407] Input: Augmented reality information

[0408] Output: The data sent.

[0409] Specific operation: The server specifies the IP address of the display device and transmits the augmented reality information.

[0410] Step 6: View AR information

[0411] The terminal (display device) displays the received augmented reality information by overlaying it on the user's field of vision using a device such as smart glasses.

[0412] Input: Data sent

[0413] Output: Augmented reality display in user's field of view

[0414] What it does: The device caches the data it receives and displays it in real time according to the user's field of view. Specifically, it overlays information about the Eiffel Tower into the user's field of view, adjusting the information in response to the user's gaze and head movements.

[0415] (Application example 1)

[0416] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal".

[0417] In conventional viewing experiences, it is difficult for users to immediately obtain information that interests them, and it is also difficult to physically experience the product, making it difficult to stimulate purchasing motivation. Furthermore, the viewing experience is often interrupted because it takes time to search and collect information. This leads to a poor user experience and a decrease in satisfaction.

[0418] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0419] In this invention, the server includes a means for image recognition of an object, a means for acquiring related information via a database or the Internet, and a means for processing the acquired related information in an AR format. This allows the user to obtain related information in real time without interrupting the viewing experience, and the product information and usage examples displayed on the smart glasses are expected to enrich the overall experience and increase purchasing motivation.

[0420] A "user" is someone who uses the smart glasses to have a viewing experience.

[0421] "Smart glasses" are glasses-type devices that have the ability to overlay augmented reality (AR) information on visual information.

[0422] A "viewing experience" is the act of a user wearing smart glasses and observing information or images.

[0423] "Real-time" refers to near-instant processing and results.

[0424] "Related information" is data that includes detailed information or use cases related to objects recognized during a viewing experience.

[0425] "AR" is an abbreviation for Augmented Reality, a technology that overlays digital information on the real world.

[0426] The term "system" refers to a collection of a series of devices and software in which a plurality of means in the present invention function in cooperation with one another.

[0427] An "object" is an object that is captured by the smart glasses' camera during the viewing experience and is the subject of image recognition.

[0428] "Image recognition" is the technology of identifying specific objects from visual data.

[0429] A "database" is a collection of information for storing and managing related information.

[0430] "Internet" is a general term for networks used to search and retrieve information.

[0431] "AR format" refers to a form of information display that uses augmented reality technology.

[0432] This invention relates to a system that provides information related to a viewing experience in real time using AR when a user is viewing an image using smart glasses. The main components and operations for specifically implementing the invention will be described below.

[0433] First, the smart glasses (terminal) are used by the user during the viewing experience and are equipped with a camera and a display. The smart glasses have the function of capturing visual information and transmitting it to a server.

[0434] The server performs image recognition of the object from the received video data. Specific image recognition algorithms such as TensorFlow and YOLO are used for this image recognition. After recognizing the object, the server retrieves related information. Related information can be retrieved from databases such as MongoDB, or the required information can be searched via the Internet.

[0435] The acquired information is converted to AR format on the server and sent to the smart glasses. The AR information is generated using AR libraries such as Vuforia and Wikitude. The smart glasses display the received AR information overlaid on the user's field of view. This allows the user to obtain relevant information in real time without interrupting the viewing experience.

[0436] The key point of the present invention is that the smart glasses can display detailed information and usage examples of products directly in the user's field of vision, enriching the overall viewing experience and increasing the user's purchasing motivation, for example, allowing a boutique customer to check the details of a garment without trying it on, or a furniture store customer to instantly learn the details of a piece of furniture.

[0437] The specific hardware and software used is as follows:

[0438] Smart glasses: Examples of smart glasses include Google Glass and Microsoft HoloLens.

[0439] Server: A cloud-based server with high-performance processing power, for example Amazon Web Services (AWS) or Microsoft Azure.

[0440] Image recognition algorithms: TensorFlow and YOLO.

[0441] Database: MongoDB or MySQL.

[0442] AR libraries: Vuforia and Wikitude.

[0443] An example of a prompt might be:

[0444] "Create an application that uses AR to display detailed information about a particular product in a virtual store."

[0445] The above is a specific embodiment of the present invention, which allows users to enrich their viewing experience and obtain relevant information in real time.

[0446] The flow of the specific process in the application example 1 will be described with reference to FIG.

[0447] Step 1:

[0448] The smart glasses (terminal) capture video data during the user's viewing experience. The camera in the smart glasses acquires the video data in real time, encodes the data in JPEG format, and sends it to the server.

[0449] Input: Video data of the user's viewing experience

[0450] Output: Encoded data in JPEG format (sent to server)

[0451] How it works: The camera in the smart glasses captures video and converts the frames into JPEG format, which is then sent to the server.

[0452] Step 2:

[0453] The server performs image recognition of the object from the received video data. It uses image recognition algorithms such as TensorFlow and YOLO to identify the object from the transmitted video data.

[0454] Input: Encoded data in JPEG format

[0455] Output: Information about the recognized object

[0456] How it works: The server receives the video data and applies image recognition algorithms to identify objects, such as clothing or furniture.

[0457] Step 3:

[0458] The server retrieves related information based on the results of image recognition, searching for and retrieving the necessary information via databases such as MongoDB or the Internet.

[0459] Input: Information about the recognized object

[0460] Output: Related information (details, use cases, historical background)

[0461] Specific operation: The server collects detailed information about the recognized object from a database or the Internet. For example, in the case of clothing, information about the material and brand is obtained.

[0462] Step 4:

[0463] The server processes the relevant information obtained into AR format, using AR libraries such as Vuforia and Wikitude to convert it into a visually easy-to-understand format.

[0464] Input: Related information (details, use cases, historical background)

[0465] Output: AR format data

[0466] Specific operation: The server generates data for AR display based on the related information collected. For example, it processes the acquired clothing information so that it can be displayed as a 3D model.

[0467] Step 5:

[0468] The server sends the processed AR information to the smart glasses, which then overlay the information onto the user's field of vision.

[0469] Input: AR format data

[0470] Output: AR information displayed on smart glasses

[0471] Specific operation: The server transmits the processed AR information to the smart glasses, which then display the information in real time, allowing users to obtain AR information that is visually intuitive.

[0472] In addition, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0473] An embodiment for implementing the present invention is a system including the following elements.

[0474] 1. Server: Obtains video frames and applies image recognition algorithms to recognize buildings and clothes worn by characters. Based on the recognition results, retrieves related information from a database and processes it in AR format. Furthermore, combines an emotion engine to recognize the user's emotions and adjusts related information and AR display. The database may be one that has been prepared in advance, or may be one that has been searched for and temporarily stored during processing using the Internet, or may be one that has been generated using a data generation model 58.

[0475] 2. Smart glasses (terminal): Receives AR information sent from the server and displays it over the user's field of view. The smart glasses recognize the user's emotions and adjust the relevant information and AR display accordingly.

[0476] As a specific example, a user wears smart glasses and watches a movie. The server retrieves frames of each scene in the movie and uses an image recognition algorithm to recognize buildings and clothes worn by characters. At the same time, it uses an emotion engine to recognize the user's emotions. For example, if the user is smiling, the emotion engine analyzes the emotion and adjusts it to provide relevant information and AR displays tailored to the user's interests and preferences.

[0477] The server processes the acquired related information in AR format and sends it to the smart glasses. The smart glasses display the received AR information, so that the user can enjoy the AR information while watching the movie. At the same time, the smart glasses recognize the user's emotions and adjust the related information and AR display according to the user's emotions. For example, if the user is smiling, the smart glasses adjust the related information and AR display to be more enjoyable.

[0478] The above is an embodiment of the present invention, and by viewing AR information provided through the smart glasses, a user can have a rich, informative and emotionally tailored video viewing experience.

[0479] The process flow will be explained below.

[0480] Step 1: The server acquires frames of a video, which allows the server to understand the scene of the video.

[0481] Step 2: The server processes the captured frames with image recognition algorithms, specifically analyzing features and patterns in the images to recognize buildings and the clothes worn by characters.

[0482] Step 3: The server retrieves related information from the database based on the recognition results, such as detailed information about the building, historical background, and the brand and styling of the clothing worn by the character.

[0483] Step 4: The server uses the emotion engine to recognize the user's emotions. Specifically, it analyzes emotions from the user's facial expressions and voice, and adjusts them to provide relevant information and AR displays that match the user's interests and preferences.

[0484] Step 5: The server combines the acquired related information with the emotion recognition results and processes the AR information. Specifically, the related information is packaged as AR data in the form of text and images, and the AR display is adjusted according to the user's emotions.

[0485] Step 6: The server sends the processed AR information to the user's smart glasses. Specifically, the server sends the AR data to the smart glasses using a communication protocol.

[0486] Step 7: The smart glasses (terminal) displays the received AR information. Specifically, by overlaying related information and AR display on the smart glasses display, the user can view the AR information while watching the video.

[0487] Step 8: The smart glasses recognize the user's emotions. Specifically, they analyze the user's emotions from their facial expressions and voice, and adjust the relevant information and AR display according to the user's emotions.

[0488] Example 2

[0489] Next, a description will be given of Example 2. In the following description, the data processing device 12 is referred to as a "server" and the smart glasses 214 are referred to as a "terminal".

[0490] In conventional video viewing systems using visual devices, the means to provide information related to the video in real time are limited, limiting the user's viewing experience. In addition, there is a lack of technology to adjust related information taking into account the user's emotions. As a result, the information obtained while watching a video may not match the user's emotional state or interests.

[0491] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0492] In this invention, the server includes a means for performing image recognition of at least a structure or clothing worn by a character in a scene of a video, a means for acquiring related information from a data storage device based on a result of the image recognition, and a means for recognizing a user's emotion using an emotion engine and adjusting the related information and the augmented reality display in order to process the acquired related information in an augmented reality format. This makes it possible to provide related information in real time in augmented reality in accordance with the user's emotional state while watching a video.

[0493] A "visual device" is a digital device that enables a user to view images and videos, including, for example, smart glasses and head-mounted displays.

[0494] "Video" refers to visual content such as television, movies, video streaming services, etc., including content that a user views through a visual device.

[0495] "Related information" refers to knowledge or data related to a particular element in a video (e.g. a building, or the clothing worn by a character, etc.), including details, historical context, styling information, etc.

[0496] "Augmented reality" refers to a technology that overlays digital information on the user's field of vision, providing real-time information to the user's field of vision.

[0497] "Image recognition" refers to the use of computer vision techniques to identify and classify objects and features within a video scene.

[0498] "Data storage device" refers to a system that stores relevant information and makes it available for search and retrieval as needed, including databases and cloud storage.

[0499] An "emotion engine" refers to algorithms and technologies that analyze human facial expressions and behavior to infer their emotional state.

[0500] "Augmented reality processing" refers to the process of adjusting and transforming acquired data to fit the user's field of view, including adding visual effects and optimizing the way it is displayed.

[0501] "Real-time" refers to the fact that processing and display occur instantly at the moment the user begins viewing.

[0502] An embodiment of this system will now be described, which mainly comprises three parties: a server, a visual device (such as smart glasses), and a user.

[0503] Hardware and Software Configuration

[0504] server

[0505] The server acts as the main processing unit and uses the following hardware and software:

[0506] Hardware: A general server machine equipped with a high-performance CPU, GPU, and large memory capacity

[0507] software:

[0508] FFmpeg: A library for capturing video frames

[0509] TensorFlow: A library for implementing image recognition algorithms

[0510] Data storage device: Database management system such as MySQL or MongoDB

[0511] Emotion engine: Emotion recognition services such as Microsoft's Emotion API

[0512] AR development tools: Unity, Vuforia

[0513] Data processing and calculation

[0514] The server first captures video data frame by frame using FFmpeg. Next, it performs image recognition on these frames using TensorFlow. Based on the recognized information on buildings and clothes, it retrieves related information from databases such as MySQL and MongoDB. After that, it analyzes the user's emotions using Emotion API and processes the related information in an augmented reality format using Unity or Vuforia.

[0515] Smart glasses (terminal)

[0516] The smart glasses have the function of receiving augmented reality information sent from a server and displaying it over the user's field of vision. This is achieved by using the following hardware and software:

[0517] Hardware: Smart glasses with built-in camera, display sensor and processor

[0518] Software: Custom application for displaying augmented reality information

[0519] User Actions

[0520] The user wears the smart glasses and watches the video. Facial expression information captured by the smart glasses while watching the video is sent to the server in real time. The server analyzes this using an emotion engine, creates augmented reality information appropriate to the user's emotional state, and sends it back to the smart glasses. While watching the video, the user can visually obtain detailed information about buildings and clothes in real time.

[0521] Examples

[0522] For example, if a user is watching a historical movie, the server will recognize a building from the movie scene and retrieve the building's historical information from the database. At the same time, if the user's facial expression indicates that they are moved, the server will send augmented reality information with an emotional effect to the smart glasses. While watching the movie, the user can enjoy information such as the background and story of the building, making the experience richer.

[0523] Examples of prompt statements

[0524] Prompt: A user is watching a movie that contains a scene depicting a certain historical building. The user's emotion is recognized as being touched. Please suggest the most suitable AR information and its presentation effect that should be displayed in this situation.

[0525] The above is an embodiment of the present invention.

[0526] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0527] The flow of this system's program processing

[0528] Step 1: Grab a video frame

[0529] The server captures frames from the video the user is watching, using the FFmpeg library.

[0530] Input: The video the user is watching

[0531] Output: Captured frame image

[0532] Specific behavior:

[0533] The server runs the FFmpeg command to capture a frame every second: ffmpeg -i input.mp4 -vf "fps=1" frame_%04d.png.

[0534] Step 2: Applying image recognition algorithms

[0535] The server uses TensorFlow to perform image recognition on the captured frames, which allows it to recognize buildings and the clothes worn by characters.

[0536] Input: The captured frame image

[0537] Output: Recognized buildings and clothing information

[0538] Specific behavior:

[0539] The server loads a TensorFlow trained model and runs each frame through the model to obtain recognition results.

[0540] Example: recognition_result = tensorflow_model.predict(frame)

[0541] Step 3: Retrieving relevant information from the database

[0542] Based on the results of image recognition, the server retrieves relevant information about buildings and clothes from a database using MySQL and MongoDB.

[0543] Input: Recognized building and clothing information

[0544] Output: Related information (details, historical background, styling information)

[0545] Specific behavior:

[0546] The server retrieves the information by executing an SQL query like SELECT FROM building_info WHERE name = 'recognized_building'.

[0547] Step 4: Emotion recognition by emotion engine

[0548] The server uses an emotion engine such as Microsoft's Emotion API to recognize the user's emotions and receives the user's facial expression data from the smart glasses.

[0549] Input: User facial expression data

[0550] Output: The perceived emotional state of the user.

[0551] Specific behavior:

[0552] The server sends the user's facial expression data to the Emotion API and obtains the emotion analysis results.

[0553] Example: emotion_result = emotion_api.analyze(user_face_image)

[0554] Step 5: Processing the AR information

[0555] The server uses Unity or Vuforia to generate and process AR information based on the acquired related information and the user's emotional state.

[0556] Input: relevant information, the user's emotional state

[0557] Output: Processed AR information

[0558] Specific behavior:

[0559] The server generates 3D objects and effects in Unity and creates the AR display.

[0560] For users who are impressed, add specific effects (eg, fireworks or music).

[0561] Step 6: Sending information to the smart glasses

[0562] The server sends the processed AR information to the smart glasses, updating the information in real time using protocols such as WebSocket.

[0563] Input: Processed AR information

[0564] Output: Sending information to smart glasses

[0565] Specific behavior:

[0566] The server sends information using websocket.send(ar_information).

[0567] Step 7: Display and emotional feedback on smart glasses

[0568] The device (smart glasses) receives the AR information from the server and displays it over the user's field of view. At the same time, it recognizes the user's emotions again and updates the information as necessary.

[0569] Input: AR information sent from the server

[0570] Output: AR information displayed in the user's field of view, and the user's emotion data obtained again.

[0571] Specific behavior:

[0572] The smart glasses display the AR information received through a display application in the field of view.

[0573] The built-in camera is used to capture the user's facial expressions in real time and the data is sent to the server.

[0574] (Application example 2)

[0575] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal".

[0576] In conventional video viewing, there was a lack of a way to provide users with detailed information about the content in real time. In addition, there was no mechanism to adjust related information according to the user's emotions, so the viewing experience was uniform and it was not possible to provide information optimized for each individual user.

[0577] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means. In this invention, the server includes a means for performing image recognition of at least a building or clothes worn by a character in a scene of a video, a means for acquiring related information from a database based on the result of the image recognition, a means for processing the acquired related information in an AR format and transmitting it to the user's smart device, a means for recognizing the user's face using a camera of the smart device and analyzing emotions, and a means for adjusting the display content of the related information based on the analyzed emotions. This makes it possible to provide the user with an optimized viewing experience by providing the user with related information in real time and further adjusting the information according to the user's emotions.

[0578] A "smart device" is a device such as a smartphone, smart glasses, or a head-mounted display that is equipped with a camera to display images and recognize the user's face.

[0579] "Video" refers to dynamic video content, such as movies, television programs, and online streaming content.

[0580] "Building" refers to any structure or structure that appears in the video.

[0581] "Clothes worn by characters" refers to the clothing worn by people or characters appearing in the video.

[0582] An "image recognition means" is an algorithm or software that extracts information from frames of video and identifies specific objects or people.

[0583] "Related information" is supplemental data provided to users, such as detailed information about buildings or clothing, historical background, design information, etc.

[0584] A "database" is a collection of related information that can be stored and retrieved as needed.

[0585] "Means for processing in AR format" refers to a technology that uses augmented reality (AR) technology to superimpose related information on the user's field of vision.

[0586] A "transmitting means" is a communication means or protocol for sending data from the server to the smart device.

[0587] "Means for recognizing faces and analyzing emotions" refers to technology that captures the user's face with a camera and determines their emotional state through facial expression analysis, etc.

[0588] "Means for adjusting display content" refers to technology for changing the content or style of information displayed in AR based on recognized emotions.

[0589] The present invention is a system that provides related information in real time based on the content of a video when a user watches the video using a smart device, and further adjusts the information according to the user's emotions. The following describes how to implement the system in detail.

[0590] First, the server captures frames from the video and uses image recognition technology to identify buildings and character clothing. Image recognition is achieved using OpenCV, a commonly used image processing library, and an AI framework equipped with a deep learning model. The server also retrieves relevant information from a database and processes it into an AR format. The database pre-stores relevant information, but it is also possible to temporarily retrieve information via the Internet during processing.

[0591] Next, the smart device (e.g., smartphone, smart glasses) receives the AR information sent from the server and displays it over the user's field of view. This uses augmented reality (AR) technology such as Google ARCore. In addition, the camera of the smart device is used to recognize the user's face and perform emotion analysis. A pre-trained emotion recognition model (e.g., a deep learning model built using Keras) is used for emotion analysis. Once the emotion is analyzed, the display content of the AR information is adjusted based on the results. For example, if the user smiles, the color and display style of the related information are adjusted to be brighter.

[0592] As a concrete example, while a user is watching their favorite movie on their smartphone, behind-the-scenes and historical information about buildings and costumes that appear in the movie are displayed in real time in AR, and the information is adjusted according to the user's emotions: if the user smiles, the information display becomes more cheerful and bright.

[0593] The specific program processing for realizing this system is as follows.

[0594] 1. The server captures frames of video and uses image recognition technology to identify buildings and character clothing.

[0595] 2. Retrieve relevant information from the database, process it into AR format and send it to the smart device.

[0596] 3. Recognize the user's face using the smart device's camera and perform emotion analysis. The emotion recognition model used is built using Keras.

[0597] 4. The display content of the AR information is adjusted based on the emotion recognition results. For example, the color and style are adjusted to be brighter when the person is smiling.

[0598] 5. The smart device will then display this tailored information, providing the user with an improved viewing experience.

[0599] An example of a prompt for a generative AI model is as follows:

[0600] Suppose a user is watching a video on their smartphone. Information about buildings and costumes that appear in the movie scenes is retrieved in real time and displayed in AR format. Furthermore, the smartphone camera recognizes the user's emotions, and if the user smiles, the display of related information is adjusted to be more enjoyable.

[0601] The flow of the specific process in the application example 2 will be described with reference to FIG.

[0602] Step 1:

[0603] The server captures the video stream from the user's smart device in real-time. It extracts frames of the video and uses them for analysis. The input is the video stream and the output is the individual frames. This is done using an image processing library like OpenCV.

[0604] Step 2:

[0605] The server applies image recognition techniques to the extracted frames. Specifically, it uses a deep learning model to identify buildings and character clothing in the frames. The input is the captured frame, and the output is the identified building and clothing information. This is done using deep learning frameworks such as TensorFlow and PyTorch.

[0606] Step 3:

[0607] The server retrieves information related to the identified buildings or clothes from a database. The input is the identified object and the output is the related information. This can be done using a SQL or NoSQL database.

[0608] Step 4:

[0609] The server processes the retrieved relevant information into an augmented reality (AR) format. The input is the relevant information, and the output is the processed information for AR. This is achieved using an AR library such as Google ARCore.

[0610] Step 5:

[0611] The server sends the processed AR information to the user's smart device. The input is the processed AR information, and the output is the data sent to the user's smart device. This is done using REST APIs and WebSockets.

[0612] Step 6:

[0613] The device uses a camera to recognize the user's face. It then uses a facial recognition algorithm to analyze the user's emotions. The input is the camera image, and the output is the user's emotional information. Keras and OpenCV are used for emotion recognition.

[0614] Step 7:

[0615] The terminal adjusts the AR information received from the server based on the user's emotion information. For example, if the user is smiling, the color and style of the displayed information is adjusted to be brighter. The input is the user's emotion information and the processed AR information, and the output is the adjusted AR information.

[0616] Step 8:

[0617] The device overlays the adjusted AR information onto the user's field of view. The input is the adjusted AR information, and the output is the AR content displayed in the user's field of view. This is done using an AR framework such as ARCore or ARKit.

[0618] As a concrete example, when a user is watching a movie, detailed information about the buildings and clothes that appear in the scene is displayed, and if the user smiles, the information is adjusted to be more pleasant.

[0619] Suppose a user is watching a video on their smartphone. Information about buildings and costumes that appear in the movie scenes is retrieved in real time and displayed in AR format. Furthermore, the smartphone camera recognizes the user's emotions, and if the user smiles, the display of related information is adjusted to be more enjoyable.

[0620] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires a voice indicating a user input for the result of the specific processing. The control unit 46A transmits the voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[0621] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by making a neural network perform deep learning. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating a voice, text data indicating a text, and image data indicating an image is input. The data generation model 58 performs inference on the input inference data according to the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0622] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0623] [Third embodiment]

[0624] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0625] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0626] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a wide area network (WAN) and / or a local area network (LAN).

[0627] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0628] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs the voice according to instructions from the processor 46.

[0629] Camera 42 is a small digital camera equipped with an optical system including a lens, an aperture, and a shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (e.g., an imaging range defined by an angle of view equivalent to the width of the field of vision of an average healthy person).

[0630] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for transmitting and receiving various types of information between the processor 46 and the processor 28 via the network 54. The transmission and reception of various types of information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is performed in a secure state.

[0631] Fig. 6 shows an example of main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0632] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32, and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0633] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0634] In the headset type terminal 314, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50, and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0635] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server", and the headset type terminal 314 will be referred to as the "terminal".

[0636] An embodiment for implementing the present invention is a system including the following elements.

[0637] 1. Server: Acquires video frames and applies image recognition algorithms to recognize buildings and clothes worn by characters. Based on the recognition results, retrieves related information from a database and processes it in AR format. The database may be one that has been prepared in advance, or may be one that has been searched for and temporarily stored during processing using the Internet, or may be one that has been generated using a data generation model 58.

[0638] 2. Smart glasses (terminal): Receives AR information sent from the server and displays it over the user's field of view. Viewers can view the AR information through the smart glasses.

[0639] In a specific example, a user wears smart glasses and watches a movie. The server retrieves frames of each scene of the movie and uses image recognition algorithms to recognize buildings and clothes worn by characters. For example, if the building appearing in the movie scene is the Eiffel Tower, the server retrieves detailed information and historical background about the Eiffel Tower from the database. Similarly, it retrieves information about the clothes worn by the characters.

[0640] The server processes the acquired related information in AR format and sends it to the smart glasses. The smart glasses display the received AR information, allowing users to enjoy the AR information while watching the movie. For example, when the Eiffel Tower appears in a scene in the movie, the smart glasses will display detailed information about the Eiffel Tower and its historical background in AR, as well as the clothes worn by the characters.

[0641] The above is an embodiment of the present invention, and by viewing AR information provided through the smart glasses, a user can have a rich and informative video viewing experience.

[0642] The process flow will be explained below.

[0643] Step 1: The server acquires frames of a video, which allows the server to understand the scene of the video.

[0644] Step 2: The server processes the captured frames with image recognition algorithms, specifically analyzing features and patterns in the images to recognize buildings and the clothes worn by characters.

[0645] Step 3: The server retrieves related information from the database based on the recognition results, such as detailed information about the building, historical background, and the brand and styling of the clothing worn by the character.

[0646] Step 4: The server processes the acquired related information in AR format and sends it to the user's smart glasses. Specifically, the server uses a communication protocol to package the related information as text and images as AR data and send it to the user's smart glasses.

[0647] Step 5: The smart glasses (terminal) displays the received AR information. Specifically, by overlaying related information on the smart glasses display, the user can view the AR information while watching the video.

[0648] Example 1

[0649] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal".

[0650] In conventional video viewing systems, it was difficult to obtain detailed information related to the video the user was watching in real time and provide it visually. In addition, there was a lack of a method to instantly provide detailed information about buildings and people's clothing while watching the video, and there were limited ways to enrich the video viewing experience.

[0651] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0652] In this invention, the server includes a means for continuously acquiring frames of a video being viewed by a display device worn by a user in order to provide information related to the video being viewed by the display device in an augmented reality format, a means for applying an image recognition algorithm to the acquired frames of the video to identify buildings and clothing of people, a means for acquiring related information related to the buildings and clothing of people from a database or a generative AI model based on the result of the image recognition, a means for converting the acquired related information into an augmented reality format, a means for transmitting the converted augmented reality information to the display device, and a means for displaying the augmented reality information by the display device in a way that is superimposed on the user's field of view. This allows the user to display detailed information acquired in real time superimposed on the video being viewed, thereby significantly improving the video viewing experience.

[0653] A "user" is a person who uses the system to watch videos and receive the augmented reality information provided.

[0654] A "display device" is a device worn by a user to display information superimposed on the user's field of vision, and includes smart glasses and the like.

[0655] "Video" refers to continuous video data such as television or movies that are viewed by users.

[0656] A "frame" is an individual still image that makes up a video.

[0657] An "image recognition algorithm" is a type of computer program that identifies objects based on captured frames.

[0658] "Buildings" refers to buildings and structures that appear in the video.

[0659] "People" refers to people who appear in the video.

[0660] "Clothing" refers to the clothes and accessories worn by a person.

[0661] "Related information" refers to detailed information or background information related to the clothing of a structure or person identified by the image recognition algorithm.

[0662] A "database" is a data storage system that stores related information and from which a server can retrieve the information.

[0663] A "generative AI model" is an artificial intelligence program that generates the necessary information by inputting a prompt sentence.

[0664] The term "augmented reality" refers to a format in which additional information is overlaid on a real-world image.

[0665] "Server" means a computer system that captures video frames, applies image recognition algorithms, captures related information, and processes and transmits the augmented reality information.

[0666] A "prompt sentence" is an input sentence that requests the generative AI model to generate specific information.

[0667] An embodiment of the present invention is a system that provides, in an augmented reality format, information related to a video being viewed by a user using a display device worn by the user.

[0668] Hardware and Software

[0669] The hardware includes the following:

[0670] High-performance GPU server (e.g. general-purpose GPU server)

[0671] Display devices (e.g. smart glasses)

[0672] The software includes the following:

[0673] OpenCV library for acquiring and processing video frames

[0674] TensorFlow Framework for Image Recognition Algorithms

[0675] Unity and ARKit / ARCore for generating augmented reality information

[0676] Database management systems for information retrieval and generative AI models (e.g., general generative AI models)

[0677] Specific explanation of operation

[0678] The server continuously captures frames of the video being viewed on the display device. For example, the server opens a movie file and uses OpenCV to extract frames every second. The server then applies a TensorFlow model to the captured frames to perform image recognition and identify buildings and people's clothing.

[0679] Relevant information about the identified structures and clothing is retrieved using a database or a generative AI model. For example, the server inputs a prompt such as "What is the history of the Eiffel Tower?" into the generative AI model and retrieves the generated information.

[0680] The acquired related information is converted into an augmented reality format using Unity or ARKit. Specifically, the generated text information is processed as a 3D object and converted into a format suitable for the display device. The converted augmented reality information is sent to the display device via Wi-Fi or Bluetooth.

[0681] The terminal (display device) displays the received augmented reality information by overlaying it on the user's field of view. For example, the display device displays information about the Eiffel Tower in the user's field of view, and when the user moves his or her head, the information moves accordingly.

[0682] Examples

[0683] Consider a scenario where a user is watching a movie through a display device worn by the user. The server captures frames of each scene of the movie every second and runs an image recognition algorithm on the captured frames. Based on the recognition results, the server retrieves detailed information about the Eiffel Tower from a database or generates information by inputting a prompt sentence such as "Tell me the history of the Eiffel Tower" into a generative AI model.

[0684] Using the prompt sentence "What is the history of the Eiffel Tower?", the AI ​​model retrieves detailed information. The information is converted into an augmented reality format and sent to the display device. The display device overlays the retrieved information on the user's field of view, allowing the user to experience the movie scene and related details in an augmented reality format.

[0685] The above system enables users to receive detailed information obtained in real time in combination with the video they are watching, greatly improving the video viewing experience.

[0686] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0687] Step 1: Grab a video frame

[0688] The server retrieves each frame sequentially from the video source: it opens the video file and uses OpenCV to extract a frame every second.

[0689] Input: Video file

[0690] Output: Still images (frames)

[0691] What happens: The server reads the stream of the video file and captures frames at regular intervals (e.g. every second).

[0692] Step 2: Applying image recognition algorithms

[0693] The server applies image recognition algorithms to the captured frames, using TensorFlow to identify buildings and people's clothing.

[0694] Input: Still image (frame)

[0695] Output: Recognition results (location and type of buildings and clothing)

[0696] Specific operation: The server runs a deep learning model (e.g., YOLO or ResNet) on the captured frames to identify objects.

[0697] Step 3: Get relevant information

[0698] Based on the recognition results, the server retrieves relevant information from a database or generative AI model, for example, detailed information about a building or a person's clothing.

[0699] Input: Recognition result

[0700] Output: Related information

[0701] Specific operation: Based on the recognition results, the server sends a query to a database to obtain the required information, or inputs a prompt sentence such as "Please tell me the history of the Eiffel Tower" into a generative AI model and obtains the generated information.

[0702] Step 4: Convert to AR format

[0703] The server converts the relevant information into an augmented reality format, using Unity or ARKit to process the information into a format suitable for the user.

[0704] Input: Related information

[0705] Output: Augmented reality information

[0706] Specific operation: The server generates the text information and images it acquires as 3D objects and animations, and processes them into an augmented reality format.

[0707] Step 5: Submit AR information

[0708] The server then sends the processed augmented reality information to the display device. Data is transferred using Wi-Fi or Bluetooth.

[0709] Input: Augmented reality information

[0710] Output: The data sent.

[0711] Specific operation: The server specifies the IP address of the display device and transmits the augmented reality information.

[0712] Step 6: View AR information

[0713] The terminal (display device) displays the received augmented reality information by overlaying it on the user's field of vision using a device such as smart glasses.

[0714] Input: Data sent

[0715] Output: Augmented reality display in user's field of view

[0716] What it does: The device caches the data it receives and displays it in real time according to the user's field of view. Specifically, it overlays information about the Eiffel Tower into the user's field of view, adjusting the information in response to the user's gaze and head movements.

[0717] (Application example 1)

[0718] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0719] In conventional viewing experiences, it is difficult for users to immediately obtain information that interests them, and it is also difficult to physically experience the product, making it difficult to stimulate purchasing motivation. Furthermore, the viewing experience is often interrupted because it takes time to search and collect information. This leads to a poor user experience and a decrease in satisfaction.

[0720] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0721] In this invention, the server includes a means for image recognition of an object, a means for acquiring related information via a database or the Internet, and a means for processing the acquired related information in an AR format. This allows the user to obtain related information in real time without interrupting the viewing experience, and the product information and usage examples displayed on the smart glasses are expected to enrich the overall experience and increase purchasing motivation.

[0722] A "user" is someone who uses the smart glasses to have a viewing experience.

[0723] "Smart glasses" are glasses-type devices that have the ability to overlay augmented reality (AR) information on visual information.

[0724] A "viewing experience" is the act of a user wearing smart glasses and observing information or images.

[0725] "Real-time" refers to near-instant processing and results.

[0726] "Related information" is data that includes detailed information or use cases related to objects recognized during a viewing experience.

[0727] "AR" is an abbreviation for Augmented Reality, a technology that overlays digital information on the real world.

[0728] The term "system" refers to a collection of a series of devices and software in which a plurality of means in the present invention function in cooperation with one another.

[0729] An "object" is an object that is captured by the smart glasses' camera during the viewing experience and is the subject of image recognition.

[0730] "Image recognition" is the technology of identifying specific objects from visual data.

[0731] A "database" is a collection of information for storing and managing related information.

[0732] "Internet" is a general term for networks used to search and retrieve information.

[0733] "AR format" refers to a form of information display that uses augmented reality technology.

[0734] This invention relates to a system that provides information related to a viewing experience in real time using AR when a user is viewing an image using smart glasses. The main components and operations for specifically implementing the invention will be described below.

[0735] First, the smart glasses (terminal) are used by the user during the viewing experience and are equipped with a camera and a display. The smart glasses have the function of capturing visual information and transmitting it to a server.

[0736] The server performs image recognition of the object from the received video data. Specific image recognition algorithms such as TensorFlow and YOLO are used for this image recognition. After recognizing the object, the server retrieves related information. Related information can be retrieved from databases such as MongoDB, or the required information can be searched via the Internet.

[0737] The acquired information is converted to AR format on the server and sent to the smart glasses. The AR information is generated using AR libraries such as Vuforia and Wikitude. The smart glasses display the received AR information overlaid on the user's field of view. This allows the user to obtain relevant information in real time without interrupting the viewing experience.

[0738] The key point of the present invention is that the smart glasses can display detailed information and usage examples of products directly in the user's field of vision, enriching the overall viewing experience and increasing the user's purchasing motivation, for example, allowing a boutique customer to check the details of a garment without trying it on, or a furniture store customer to instantly learn the details of a piece of furniture.

[0739] The specific hardware and software used is as follows:

[0740] Smart glasses: Examples of smart glasses include Google Glass and Microsoft HoloLens.

[0741] Server: A cloud-based server with high-performance processing power, for example Amazon Web Services (AWS) or Microsoft Azure.

[0742] Image recognition algorithms: TensorFlow and YOLO.

[0743] Database: MongoDB or MySQL.

[0744] AR libraries: Vuforia and Wikitude.

[0745] An example of a prompt might be:

[0746] "Create an application that uses AR to display detailed information about a particular product in a virtual store."

[0747] The above is a specific embodiment of the present invention, which allows users to enrich their viewing experience and obtain relevant information in real time.

[0748] The flow of the specific process in the application example 1 will be described with reference to FIG.

[0749] Step 1:

[0750] The smart glasses (terminal) capture video data during the user's viewing experience. The camera in the smart glasses acquires the video data in real time, encodes the data in JPEG format, and sends it to the server.

[0751] Input: Video data of the user's viewing experience

[0752] Output: Encoded data in JPEG format (sent to server)

[0753] How it works: The camera in the smart glasses captures video and converts the frames into JPEG format, which is then sent to the server.

[0754] Step 2:

[0755] The server performs image recognition of the object from the received video data. It uses image recognition algorithms such as TensorFlow and YOLO to identify the object from the transmitted video data.

[0756] Input: Encoded data in JPEG format

[0757] Output: Information about the recognized object

[0758] How it works: The server receives the video data and applies image recognition algorithms to identify objects, such as clothing or furniture.

[0759] Step 3:

[0760] The server retrieves related information based on the results of image recognition, searching for and retrieving the necessary information via databases such as MongoDB or the Internet.

[0761] Input: Information about the recognized object

[0762] Output: Related information (details, use cases, historical background)

[0763] Specific operation: The server collects detailed information about the recognized object from a database or the Internet. For example, in the case of clothing, information about the material and brand is obtained.

[0764] Step 4:

[0765] The server processes the relevant information obtained into AR format, using AR libraries such as Vuforia and Wikitude to convert it into a visually easy-to-understand format.

[0766] Input: Related information (details, use cases, historical background)

[0767] Output: AR format data

[0768] Specific operation: The server generates data for AR display based on the related information collected. For example, it processes the acquired clothing information so that it can be displayed as a 3D model.

[0769] Step 5:

[0770] The server sends the processed AR information to the smart glasses, which then overlay the information onto the user's field of vision.

[0771] Input: AR format data

[0772] Output: AR information displayed on smart glasses

[0773] Specific operation: The server transmits the processed AR information to the smart glasses, which then display the information in real time, allowing users to obtain AR information that is visually intuitive.

[0774] In addition, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0775] An embodiment for implementing the present invention is a system including the following elements.

[0776] 1. Server: Obtains video frames and applies image recognition algorithms to recognize buildings and clothes worn by characters. Based on the recognition results, retrieves related information from a database and processes it in AR format. Furthermore, combines an emotion engine to recognize the user's emotions and adjusts related information and AR display. The database may be one that has been prepared in advance, or may be one that has been searched for and temporarily stored during processing using the Internet, or may be one that has been generated using a data generation model 58.

[0777] 2. Smart glasses (terminal): Receives AR information sent from the server and displays it over the user's field of view. The smart glasses recognize the user's emotions and adjust the relevant information and AR display accordingly.

[0778] As a specific example, a user wears smart glasses and watches a movie. The server retrieves frames of each scene in the movie and uses an image recognition algorithm to recognize buildings and clothes worn by characters. At the same time, it uses an emotion engine to recognize the user's emotions. For example, if the user is smiling, the emotion engine analyzes the emotion and adjusts it to provide relevant information and AR displays tailored to the user's interests and preferences.

[0779] The server processes the acquired related information in AR format and sends it to the smart glasses. The smart glasses display the received AR information, so that the user can enjoy the AR information while watching the movie. At the same time, the smart glasses recognize the user's emotions and adjust the related information and AR display according to the user's emotions. For example, if the user is smiling, the smart glasses adjust the related information and AR display to be more enjoyable.

[0780] The above is an embodiment of the present invention, and by viewing AR information provided through the smart glasses, a user can have a rich, informative and emotionally tailored video viewing experience.

[0781] The process flow will be explained below.

[0782] Step 1: The server acquires frames of a video, which allows the server to understand the scene of the video.

[0783] Step 2: The server processes the captured frames with image recognition algorithms, specifically analyzing features and patterns in the images to recognize buildings and the clothes worn by characters.

[0784] Step 3: The server retrieves related information from the database based on the recognition results, such as detailed information about the building, historical background, and the brand and styling of the clothing worn by the character.

[0785] Step 4: The server uses the emotion engine to recognize the user's emotions. Specifically, it analyzes emotions from the user's facial expressions and voice, and adjusts them to provide relevant information and AR displays that match the user's interests and preferences.

[0786] Step 5: The server combines the acquired related information with the emotion recognition results and processes the AR information. Specifically, the related information is packaged as AR data in the form of text and images, and the AR display is adjusted according to the user's emotions.

[0787] Step 6: The server sends the processed AR information to the user's smart glasses. Specifically, the server sends the AR data to the smart glasses using a communication protocol.

[0788] Step 7: The smart glasses (terminal) displays the received AR information. Specifically, by overlaying related information and AR display on the smart glasses display, the user can view the AR information while watching the video.

[0789] Step 8: The smart glasses recognize the user's emotions. Specifically, they analyze the user's emotions from their facial expressions and voice, and adjust the relevant information and AR display according to the user's emotions.

[0790] Example 2

[0791] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal".

[0792] In conventional video viewing systems using visual devices, the means to provide information related to the video in real time are limited, limiting the user's viewing experience. In addition, there is a lack of technology to adjust related information taking into account the user's emotions. As a result, the information obtained while watching a video may not match the user's emotional state or interests.

[0793] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0794] In this invention, the server includes a means for performing image recognition of at least a structure or clothing worn by a character in a scene of a video, a means for acquiring related information from a data storage device based on a result of the image recognition, and a means for recognizing a user's emotion using an emotion engine and adjusting the related information and the augmented reality display in order to process the acquired related information in an augmented reality format. This makes it possible to provide related information in real time in augmented reality in accordance with the user's emotional state while watching a video.

[0795] A "visual device" is a digital device that enables a user to view images and videos, including, for example, smart glasses and head-mounted displays.

[0796] "Video" refers to visual content such as television, movies, video streaming services, etc., including content that a user views through a visual device.

[0797] "Related information" refers to knowledge or data related to a particular element in a video (e.g. a building, or the clothing worn by a character), including details, historical context, styling information, etc.

[0798] "Augmented reality" refers to a technology that overlays digital information on the user's field of vision, providing real-time information to the user's field of vision.

[0799] "Image recognition" refers to the use of computer vision techniques to identify and classify objects and features within a video scene.

[0800] "Data storage device" refers to a system that stores relevant information and makes it available for search and retrieval as needed, including databases and cloud storage.

[0801] An "emotion engine" refers to algorithms and technologies that analyze human facial expressions and behavior to infer their emotional state.

[0802] "Augmented reality processing" refers to the process of adjusting and transforming acquired data to fit the user's field of view, including adding visual effects and optimizing the way it is displayed.

[0803] "Real-time" refers to the fact that processing and display occur instantly at the moment the user begins viewing.

[0804] An embodiment of this system will now be described, which mainly comprises three parties: a server, a visual device (such as smart glasses), and a user.

[0805] Hardware and Software Configuration

[0806] server

[0807] The server acts as the main processing unit and uses the following hardware and software:

[0808] Hardware: A general server machine equipped with a high-performance CPU, GPU, and large memory capacity

[0809] software:

[0810] FFmpeg: A library for capturing video frames

[0811] TensorFlow: A library for implementing image recognition algorithms

[0812] Data storage device: Database management system such as MySQL or MongoDB

[0813] Emotion engine: Emotion recognition services such as Microsoft's Emotion API

[0814] AR development tools: Unity, Vuforia

[0815] Data processing and calculation

[0816] The server first captures video data frame by frame using FFmpeg. Next, it performs image recognition on these frames using TensorFlow. Based on the recognized information on buildings and clothes, it retrieves related information from databases such as MySQL and MongoDB. After that, it analyzes the user's emotions using Emotion API and processes the related information in an augmented reality format using Unity or Vuforia.

[0817] Smart glasses (terminal)

[0818] The smart glasses have the function of receiving augmented reality information sent from a server and displaying it over the user's field of vision. This is achieved by using the following hardware and software:

[0819] Hardware: Smart glasses with built-in camera, display sensor and processor

[0820] Software: Custom application for displaying augmented reality information

[0821] User Actions

[0822] The user wears the smart glasses and watches the video. Facial expression information captured by the smart glasses while watching the video is sent to the server in real time. The server analyzes this using an emotion engine, creates augmented reality information appropriate to the user's emotional state, and sends it back to the smart glasses. While watching the video, the user can visually obtain detailed information about buildings and clothes in real time.

[0823] Examples

[0824] For example, if a user is watching a historical movie, the server will recognize a building from the movie scene and retrieve the building's historical information from the database. At the same time, if the user's facial expression indicates that they are moved, the server will send augmented reality information with an emotional effect to the smart glasses. While watching the movie, the user can enjoy information such as the background and story of the building, making the experience richer.

[0825] Examples of prompt statements

[0826] Prompt: A user is watching a movie that contains a scene depicting a certain historical building. The user's emotion is recognized as being touched. Please suggest the most suitable AR information and its presentation effect that should be displayed in this situation.

[0827] The above is an embodiment of the present invention.

[0828] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0829] The flow of this system's program processing

[0830] Step 1: Grab a video frame

[0831] The server captures frames from the video the user is watching, using the FFmpeg library.

[0832] Input: The video the user is watching

[0833] Output: Captured frame image

[0834] Specific behavior:

[0835] The server runs the FFmpeg command to capture a frame every second: ffmpeg -i input.mp4 -vf "fps=1" frame_%04d.png.

[0836] Step 2: Applying image recognition algorithms

[0837] The server uses TensorFlow to perform image recognition on the captured frames, which allows it to recognize buildings and the clothes worn by characters.

[0838] Input: The captured frame image

[0839] Output: Recognized buildings and clothing information

[0840] Specific behavior:

[0841] The server loads a TensorFlow trained model and runs each frame through the model to obtain recognition results.

[0842] Example: recognition_result = tensorflow_model.predict(frame)

[0843] Step 3: Retrieving relevant information from the database

[0844] Based on the results of image recognition, the server retrieves relevant information about buildings and clothes from a database using MySQL and MongoDB.

[0845] Input: Recognized building and clothing information

[0846] Output: Related information (details, historical background, styling information)

[0847] Specific behavior:

[0848] The server retrieves the information by executing an SQL query like SELECT FROM building_info WHERE name = 'recognized_building'.

[0849] Step 4: Emotion recognition by emotion engine

[0850] The server uses an emotion engine such as Microsoft's Emotion API to recognize the user's emotions and receives the user's facial expression data from the smart glasses.

[0851] Input: User facial expression data

[0852] Output: The perceived emotional state of the user.

[0853] Specific behavior:

[0854] The server sends the user's facial expression data to the Emotion API and obtains the emotion analysis results.

[0855] Example: emotion_result = emotion_api.analyze(user_face_image)

[0856] Step 5: Processing the AR information

[0857] The server uses Unity or Vuforia to generate and process AR information based on the acquired related information and the user's emotional state.

[0858] Input: relevant information, the user's emotional state

[0859] Output: Processed AR information

[0860] Specific behavior:

[0861] The server generates 3D objects and effects in Unity and creates the AR display.

[0862] For users who are impressed, add specific effects (eg, fireworks or music).

[0863] Step 6: Sending information to the smart glasses

[0864] The server sends the processed AR information to the smart glasses, updating the information in real time using protocols such as WebSocket.

[0865] Input: Processed AR information

[0866] Output: Sending information to smart glasses

[0867] Specific behavior:

[0868] The server sends information using websocket.send(ar_information).

[0869] Step 7: Display and emotional feedback on smart glasses

[0870] The device (smart glasses) receives the AR information from the server and displays it over the user's field of view. At the same time, it recognizes the user's emotions again and updates the information as necessary.

[0871] Input: AR information sent from the server

[0872] Output: AR information displayed in the user's field of view, and the user's emotion data obtained again.

[0873] Specific behavior:

[0874] The smart glasses display the AR information received through a display application in the field of view.

[0875] The built-in camera is used to capture the user's facial expressions in real time and the data is sent to the server.

[0876] (Application example 2)

[0877] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server", and the headset type terminal 314 will be referred to as a "terminal".

[0878] In conventional video viewing, there was a lack of a way to provide users with detailed information about the content in real time. In addition, there was no mechanism to adjust related information according to the user's emotions, so the viewing experience was uniform and it was not possible to provide information optimized for each individual user.

[0879] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means. In this invention, the server includes a means for performing image recognition of at least a building or clothes worn by a character in a scene of a video, a means for acquiring related information from a database based on the result of the image recognition, a means for processing the acquired related information in an AR format and transmitting it to the user's smart device, a means for recognizing the user's face using a camera of the smart device and analyzing emotions, and a means for adjusting the display content of the related information based on the analyzed emotions. This makes it possible to provide the user with an optimized viewing experience by providing the user with related information in real time and further adjusting the information according to the user's emotions.

[0880] A "smart device" is a device such as a smartphone, smart glasses, or a head-mounted display that is equipped with a camera to display images and recognize the user's face.

[0881] "Video" refers to dynamic video content, such as movies, television programs, and online streaming content.

[0882] "Building" refers to any structure or structure that appears in the video.

[0883] "Clothes worn by characters" refers to the clothing worn by people or characters appearing in the video.

[0884] An "image recognition means" is an algorithm or software that extracts information from frames of video and identifies specific objects or people.

[0885] "Related information" is supplemental data provided to users, such as detailed information about buildings or clothing, historical background, design information, etc.

[0886] A "database" is a collection of related information that can be stored and retrieved as needed.

[0887] "Means for processing in AR format" refers to a technology that uses augmented reality (AR) technology to superimpose related information on the user's field of vision.

[0888] A "transmitting means" is a communication means or protocol for sending data from the server to the smart device.

[0889] "Means for recognizing faces and analyzing emotions" refers to technology that captures the user's face with a camera and determines their emotional state through facial expression analysis, etc.

[0890] "Means for adjusting display content" refers to technology for changing the content or style of information displayed in AR based on recognized emotions.

[0891] The present invention is a system that provides related information in real time based on the content of a video when a user watches the video using a smart device, and further adjusts the information according to the user's emotions. The following describes how to implement the system in detail.

[0892] First, the server captures frames from the video and uses image recognition technology to identify buildings and character clothing. Image recognition is achieved using OpenCV, a commonly used image processing library, and an AI framework equipped with a deep learning model. The server also retrieves relevant information from a database and processes it into an AR format. The database pre-stores relevant information, but it is also possible to temporarily retrieve information via the Internet during processing.

[0893] Next, the smart device (e.g., smartphone, smart glasses) receives the AR information sent from the server and displays it over the user's field of view. This uses augmented reality (AR) technology such as Google ARCore. In addition, the camera of the smart device is used to recognize the user's face and perform emotion analysis. A pre-trained emotion recognition model (e.g., a deep learning model built using Keras) is used for emotion analysis. Once the emotion is analyzed, the display content of the AR information is adjusted based on the results. For example, if the user smiles, the color and display style of the related information are adjusted to be brighter.

[0894] As a concrete example, while a user is watching their favorite movie on their smartphone, behind-the-scenes and historical information about buildings and costumes that appear in the movie are displayed in real time in AR, and the information is adjusted according to the user's emotions: if the user smiles, the information display becomes more cheerful and bright.

[0895] The specific program processing for realizing this system is as follows.

[0896] 1. The server captures frames of video and uses image recognition technology to identify buildings and character clothing.

[0897] 2. Retrieve relevant information from the database, process it into AR format and send it to the smart device.

[0898] 3. Recognize the user's face using the smart device's camera and perform emotion analysis. The emotion recognition model used is built using Keras.

[0899] 4. The display content of the AR information is adjusted based on the emotion recognition results. For example, the color and style are adjusted to be brighter when the person is smiling.

[0900] 5. The smart device will then display this tailored information, providing the user with an improved viewing experience.

[0901] An example of a prompt for a generative AI model is as follows:

[0902] Suppose a user is watching a video on their smartphone. Information about buildings and costumes that appear in the movie scenes is retrieved in real time and displayed in AR format. Furthermore, the smartphone camera recognizes the user's emotions, and if the user smiles, the display of related information is adjusted to be more enjoyable.

[0903] The flow of the specific process in the application example 2 will be described with reference to FIG.

[0904] Step 1:

[0905] The server captures the video stream from the user's smart device in real-time. It extracts frames of the video and uses them for analysis. The input is the video stream and the output are the individual frames. This is done using an image processing library like OpenCV.

[0906] Step 2:

[0907] The server applies image recognition techniques to the extracted frames. Specifically, it uses a deep learning model to identify buildings and character clothing in the frames. The input is the captured frame, and the output is the identified building and clothing information. This is done using deep learning frameworks such as TensorFlow and PyTorch.

[0908] Step 3:

[0909] The server retrieves information related to the identified buildings or clothes from a database. The input is the identified object and the output is the related information. This can be done using a SQL or NoSQL database.

[0910] Step 4:

[0911] The server processes the retrieved relevant information into an augmented reality (AR) format. The input is the relevant information, and the output is the processed information for AR. This is achieved using an AR library such as Google ARCore.

[0912] Step 5:

[0913] The server sends the processed AR information to the user's smart device. The input is the processed AR information, and the output is the data sent to the user's smart device. This is done using REST APIs and WebSockets.

[0914] Step 6:

[0915] The device uses a camera to recognize the user's face. It then uses a facial recognition algorithm to analyze the user's emotions. The input is the camera image, and the output is the user's emotional information. Keras and OpenCV are used for emotion recognition.

[0916] Step 7:

[0917] The terminal adjusts the AR information received from the server based on the user's emotion information. For example, if the user is smiling, the color and style of the displayed information is adjusted to be brighter. The input is the user's emotion information and the processed AR information, and the output is the adjusted AR information.

[0918] Step 8:

[0919] The device displays the adjusted AR information overlaid on the user's field of view. The input is the adjusted AR information, and the output is the AR content displayed in the user's field of view. This is done using an AR framework such as ARCore or ARKit.

[0920] As a concrete example, when a user is watching a movie, detailed information about the buildings and clothes that appear in the scene is displayed, and if the user smiles, the information is adjusted to be more pleasant.

[0921] Suppose a user is watching a video on their smartphone. Information about buildings and costumes that appear in the movie scenes is retrieved in real time and displayed in AR format. Furthermore, the smartphone camera recognizes the user's emotions, and if the user smiles, the display of related information is adjusted to be more enjoyable.

[0922] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input for the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[0923] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by making a neural network perform deep learning. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating a voice, text data indicating a text, and image data indicating an image is input. The data generation model 58 performs inference on the input inference data according to the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0924] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[0925] [Fourth embodiment]

[0926] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[0927] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0928] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a wide area network (WAN) and / or a local area network (LAN).

[0929] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. In addition, the microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[0930] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs the voice according to instructions from the processor 46.

[0931] Camera 42 is a small digital camera equipped with an optical system including a lens, an aperture, and a shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (e.g., an imaging range defined by an angle of view equivalent to the width of the field of vision of an average healthy person).

[0932] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for transmitting and receiving various types of information between the processor 46 and the processor 28 via the network 54. The transmission and reception of various types of information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is performed in a secure state.

[0933] The control target 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, legs, etc. The posture and behavior of the robot 414 are controlled by controlling the motors of the arms, hands, legs, etc. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[0934] Fig. 8 shows an example of main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0935] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32, and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0936] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0937] In the robot 414, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50, and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0938] Next, a description will be given of the specific processing by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal".

[0939] An embodiment for implementing the present invention is a system including the following elements.

[0940] 1. Server: Acquires video frames and applies image recognition algorithms to recognize buildings and clothes worn by characters. Based on the recognition results, retrieves related information from a database and processes it in AR format. The database may be one that has been prepared in advance, or may be one that has been searched for and temporarily stored during processing using the Internet, or may be one that has been generated using a data generation model 58.

[0941] 2. Smart glasses (terminal): Receives AR information sent from the server and displays it over the user's field of view. Viewers can view the AR information through the smart glasses.

[0942] In a specific example, a user wears smart glasses and watches a movie. The server retrieves frames of each scene of the movie and uses image recognition algorithms to recognize buildings and clothes worn by characters. For example, if the building appearing in the movie scene is the Eiffel Tower, the server retrieves detailed information and historical background about the Eiffel Tower from the database. Similarly, it retrieves information about the clothes worn by the characters.

[0943] The server processes the acquired related information in AR format and sends it to the smart glasses. The smart glasses display the received AR information, allowing users to enjoy the AR information while watching the movie. For example, when the Eiffel Tower appears in a scene in the movie, the smart glasses will display detailed information about the Eiffel Tower and its historical background in AR, as well as the clothes worn by the characters.

[0944] The above is an embodiment of the present invention, and by viewing AR information provided through the smart glasses, a user can have a rich and informative video viewing experience.

[0945] The process flow will be explained below.

[0946] Step 1: The server acquires frames of a video, which allows the server to understand the scene of the video.

[0947] Step 2: The server processes the captured frames with image recognition algorithms, specifically analyzing features and patterns in the images to recognize buildings and the clothes worn by characters.

[0948] Step 3: The server retrieves related information from the database based on the recognition results, such as detailed information about the building, historical background, and the brand and styling of the clothing worn by the character.

[0949] Step 4: The server processes the acquired related information in AR format and sends it to the user's smart glasses. Specifically, the server uses a communication protocol to package the related information as text and images as AR data and send it to the user's smart glasses.

[0950] Step 5: The smart glasses (terminal) displays the received AR information. Specifically, by overlaying related information on the smart glasses display, the user can view the AR information while watching the video.

[0951] Example 1

[0952] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal".

[0953] In conventional video viewing systems, it was difficult to obtain detailed information related to the video the user was watching in real time and provide it visually. In addition, there was a lack of a method to instantly provide detailed information about buildings and people's clothing while watching the video, and there were limited ways to enrich the video viewing experience.

[0954] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0955] In this invention, the server includes a means for continuously acquiring frames of a video being viewed by a display device worn by a user in order to provide information related to the video being viewed by the display device in an augmented reality format, a means for applying an image recognition algorithm to the acquired frames of the video to identify buildings and clothing of people, a means for acquiring related information related to the buildings and clothing of people from a database or a generative AI model based on the result of the image recognition, a means for converting the acquired related information into an augmented reality format, a means for transmitting the converted augmented reality information to the display device, and a means for displaying the augmented reality information by the display device in a way that is superimposed on the user's field of view. This allows the user to display detailed information acquired in real time superimposed on the video being viewed, thereby significantly improving the video viewing experience.

[0956] A "user" is a person who uses the system to watch videos and receive the augmented reality information provided.

[0957] A "display device" is a device worn by a user to display information superimposed on the user's field of vision, and includes smart glasses and the like.

[0958] "Video" refers to continuous video data such as television or movies that are viewed by users.

[0959] A "frame" is an individual still image that makes up a video.

[0960] An "image recognition algorithm" is a type of computer program that identifies objects based on captured frames.

[0961] "Buildings" refers to buildings and structures that appear in the video.

[0962] "People" refers to people who appear in the video.

[0963] "Clothing" refers to the clothes and accessories worn by a person.

[0964] "Related information" refers to detailed information or background information related to the clothing of a structure or person identified by the image recognition algorithm.

[0965] A "database" is a data storage system that stores related information and from which a server can retrieve the information.

[0966] A "generative AI model" is an artificial intelligence program that generates the necessary information by inputting a prompt sentence.

[0967] The term "augmented reality" refers to a format in which additional information is overlaid on a real-world image.

[0968] "Server" means a computer system that captures video frames, applies image recognition algorithms, captures related information, and processes and transmits the augmented reality information.

[0969] A "prompt sentence" is an input sentence that requests the generative AI model to generate specific information.

[0970] An embodiment of the present invention is a system that provides, in an augmented reality format, information related to a video being viewed by a user using a display device worn by the user.

[0971] Hardware and Software

[0972] The hardware includes the following:

[0973] High-performance GPU server (e.g. general-purpose GPU server)

[0974] Display devices (e.g. smart glasses)

[0975] The software includes the following:

[0976] OpenCV library for acquiring and processing video frames

[0977] TensorFlow Framework for Image Recognition Algorithms

[0978] Unity and ARKit / ARCore for generating augmented reality information

[0979] Database management systems for information retrieval and generative AI models (e.g., general generative AI models)

[0980] Specific explanation of operation

[0981] The server continuously captures frames of the video being viewed on the display device. For example, the server opens a movie file and uses OpenCV to extract frames every second. The server then applies a TensorFlow model to the captured frames to perform image recognition and identify buildings and people's clothing.

[0982] Relevant information about the identified structures and clothing is retrieved using a database or a generative AI model. For example, the server inputs a prompt such as "What is the history of the Eiffel Tower?" into the generative AI model and retrieves the generated information.

[0983] The acquired related information is converted into an augmented reality format using Unity or ARKit. Specifically, the generated text information is processed as a 3D object and converted into a format suitable for the display device. The converted augmented reality information is sent to the display device via Wi-Fi or Bluetooth.

[0984] The terminal (display device) displays the received augmented reality information by overlaying it on the user's field of view. For example, the display device displays information about the Eiffel Tower in the user's field of view, and when the user moves his or her head, the information moves accordingly.

[0985] Examples

[0986] Consider a scenario where a user is watching a movie through a display device worn by the user. The server captures frames of each scene of the movie every second and runs an image recognition algorithm on the captured frames. Based on the recognition results, the server retrieves detailed information about the Eiffel Tower from a database or generates information by inputting a prompt sentence such as "Tell me the history of the Eiffel Tower" into a generative AI model.

[0987] Using the prompt sentence "What is the history of the Eiffel Tower?", the AI ​​model retrieves detailed information. The information is converted into an augmented reality format and sent to the display device. The display device overlays the retrieved information on the user's field of view, allowing the user to experience the movie scene and related details in an augmented reality format.

[0988] The above system enables users to receive detailed information obtained in real time in combination with the video they are watching, greatly improving the video viewing experience.

[0989] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0990] Step 1: Grab a video frame

[0991] The server retrieves each frame sequentially from the video source: it opens the video file and uses OpenCV to extract a frame every second.

[0992] Input: Video file

[0993] Output: Still images (frames)

[0994] What happens: The server reads the stream of the video file and captures frames at regular intervals (e.g. every second).

[0995] Step 2: Applying image recognition algorithms

[0996] The server applies image recognition algorithms to the captured frames, using TensorFlow to identify buildings and people's clothing.

[0997] Input: Still image (frame)

[0998] Output: Recognition results (location and type of buildings and clothing)

[0999] Specific operation: The server runs a deep learning model (e.g., YOLO or ResNet) on the captured frames to identify objects.

[1000] Step 3: Get relevant information

[1001] Based on the recognition results, the server retrieves relevant information from a database or generative AI model, for example, detailed information about a building or a person's clothing.

[1002] Input: Recognition result

[1003] Output: Related information

[1004] Specific operation: Based on the recognition results, the server sends a query to a database to obtain the required information, or inputs a prompt sentence such as "Please tell me the history of the Eiffel Tower" into a generative AI model and obtains the generated information.

[1005] Step 4: Convert to AR format

[1006] The server converts the relevant information into an augmented reality format, using Unity or ARKit to process the information into a format suitable for the user.

[1007] Input: Related information

[1008] Output: Augmented reality information

[1009] Specific operation: The server generates the text information and images it acquires as 3D objects and animations, and processes them into an augmented reality format.

[1010] Step 5: Submit AR information

[1011] The server then sends the processed augmented reality information to the display device. Data is transferred using Wi-Fi or Bluetooth.

[1012] Input: Augmented reality information

[1013] Output: The data sent.

[1014] Specific operation: The server specifies the IP address of the display device and transmits the augmented reality information.

[1015] Step 6: View AR information

[1016] The terminal (display device) displays the received augmented reality information by overlaying it on the user's field of vision using a device such as smart glasses.

[1017] Input: Data sent

[1018] Output: Augmented reality display in user's field of view

[1019] What it does: The device caches the data it receives and displays it in real time according to the user's field of view. Specifically, it overlays information about the Eiffel Tower into the user's field of view, adjusting the information in response to the user's gaze and head movements.

[1020] (Application example 1)

[1021] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1022] In conventional viewing experiences, it is difficult for users to immediately obtain information that interests them, and it is also difficult to physically experience the product, making it difficult to stimulate purchasing motivation. Furthermore, the viewing experience is often interrupted because it takes time to search and collect information. This leads to a poor user experience and a decrease in satisfaction.

[1023] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1024] In this invention, the server includes a means for image recognition of an object, a means for acquiring related information via a database or the Internet, and a means for processing the acquired related information in an AR format. This allows the user to obtain related information in real time without interrupting the viewing experience, and the product information and usage examples displayed on the smart glasses are expected to enrich the overall experience and increase purchasing motivation.

[1025] A "user" is someone who uses the smart glasses to have a viewing experience.

[1026] "Smart glasses" are glasses-type devices that have the ability to overlay augmented reality (AR) information on visual information.

[1027] A "viewing experience" is the act of a user wearing smart glasses and observing information or images.

[1028] "Real-time" refers to near-instant processing and results.

[1029] "Related information" is data that includes detailed information or use cases related to objects recognized during a viewing experience.

[1030] "AR" is an abbreviation for Augmented Reality, a technology that overlays digital information on the real world.

[1031] The term "system" refers to a collection of a series of devices and software in which a plurality of means in the present invention function in cooperation with one another.

[1032] An "object" is an object that is captured by the smart glasses' camera during the viewing experience and is the subject of image recognition.

[1033] "Image recognition" is the technology of identifying specific objects from visual data.

[1034] A "database" is a collection of information for storing and managing related information.

[1035] "Internet" is a general term for networks used to search and retrieve information.

[1036] "AR format" refers to a form of information display that uses augmented reality technology.

[1037] This invention relates to a system that provides information related to a viewing experience in real time using AR when a user is viewing an image using smart glasses. The main components and operations for specifically implementing the invention will be described below.

[1038] First, the smart glasses (terminal) are used by the user during the viewing experience and are equipped with a camera and a display. The smart glasses have the function of capturing visual information and transmitting it to a server.

[1039] The server performs image recognition of the object from the received video data. Specific image recognition algorithms such as TensorFlow and YOLO are used for this image recognition. After recognizing the object, the server retrieves related information. Related information can be retrieved from databases such as MongoDB, or the required information can be searched via the Internet.

[1040] The acquired information is converted to AR format on the server and sent to the smart glasses. The AR information is generated using AR libraries such as Vuforia and Wikitude. The smart glasses display the received AR information overlaid on the user's field of view. This allows the user to obtain relevant information in real time without interrupting the viewing experience.

[1041] The key point of the present invention is that the smart glasses can display detailed information and usage examples of products directly in the user's field of vision, enriching the overall viewing experience and increasing the user's purchasing motivation, for example, allowing a boutique customer to check the details of a garment without trying it on, or a furniture store customer to instantly learn the details of a piece of furniture.

[1042] The specific hardware and software used is as follows:

[1043] Smart glasses: Examples of smart glasses include Google Glass and Microsoft HoloLens.

[1044] Server: A cloud-based server with high-performance processing power, for example Amazon Web Services (AWS) or Microsoft Azure.

[1045] Image recognition algorithms: TensorFlow and YOLO.

[1046] Database: MongoDB or MySQL.

[1047] AR libraries: Vuforia and Wikitude.

[1048] An example of a prompt might be:

[1049] "Create an application that uses AR to display detailed information about a particular product in a virtual store."

[1050] The above is a specific embodiment of the present invention, which allows users to enrich their viewing experience and obtain relevant information in real time.

[1051] The flow of the specific process in the application example 1 will be described with reference to FIG.

[1052] Step 1:

[1053] The smart glasses (terminal) capture video data during the user's viewing experience. The camera in the smart glasses acquires the video data in real time, encodes the data in JPEG format, and sends it to the server.

[1054] Input: Video data of the user's viewing experience

[1055] Output: Encoded data in JPEG format (sent to server)

[1056] How it works: The camera in the smart glasses captures video and converts the frames into JPEG format, which is then sent to the server.

[1057] Step 2:

[1058] The server performs image recognition of the object from the received video data. It uses image recognition algorithms such as TensorFlow and YOLO to identify the object from the transmitted video data.

[1059] Input: Encoded data in JPEG format

[1060] Output: Information about the recognized object

[1061] How it works: The server receives the video data and applies image recognition algorithms to identify objects, such as clothing or furniture.

[1062] Step 3:

[1063] The server retrieves related information based on the results of image recognition, searching for and retrieving the necessary information via databases such as MongoDB or the Internet.

[1064] Input: Information about the recognized object

[1065] Output: Related information (details, use cases, historical background)

[1066] Specific operation: The server collects detailed information about the recognized object from a database or the Internet. For example, in the case of clothing, information about the material and brand is obtained.

[1067] Step 4:

[1068] The server processes the relevant information obtained into AR format, using AR libraries such as Vuforia and Wikitude to convert it into a visually easy-to-understand format.

[1069] Input: Related information (details, use cases, historical background)

[1070] Output: AR format data

[1071] Specific operation: The server generates data for AR display based on the related information collected. For example, it processes the acquired clothing information so that it can be displayed as a 3D model.

[1072] Step 5:

[1073] The server sends the processed AR information to the smart glasses, which then overlay the information onto the user's field of vision.

[1074] Input: AR format data

[1075] Output: AR information displayed on smart glasses

[1076] Specific operation: The server transmits the processed AR information to the smart glasses, which then display the information in real time, allowing users to obtain AR information that is visually intuitive.

[1077] In addition, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1078] An embodiment for implementing the present invention is a system including the following elements.

[1079] 1. Server: Obtains video frames and applies image recognition algorithms to recognize buildings and clothes worn by characters. Based on the recognition results, retrieves related information from a database and processes it in AR format. Furthermore, combines an emotion engine to recognize the user's emotions and adjusts related information and AR display. The database may be one that has been prepared in advance, or may be one that has been searched for and temporarily stored during processing using the Internet, or may be one that has been generated using a data generation model 58.

[1080] 2. Smart glasses (terminal): Receives AR information sent from the server and displays it over the user's field of view. The smart glasses recognize the user's emotions and adjust the relevant information and AR display accordingly.

[1081] As a specific example, a user wears smart glasses and watches a movie. The server retrieves frames of each scene in the movie and uses an image recognition algorithm to recognize buildings and clothes worn by characters. At the same time, it uses an emotion engine to recognize the user's emotions. For example, if the user is smiling, the emotion engine analyzes the emotion and adjusts it to provide relevant information and AR displays tailored to the user's interests and preferences.

[1082] The server processes the acquired related information in AR format and sends it to the smart glasses. The smart glasses display the received AR information, so that the user can enjoy the AR information while watching the movie. At the same time, the smart glasses recognize the user's emotions and adjust the related information and AR display according to the user's emotions. For example, if the user is smiling, the smart glasses adjust the related information and AR display to be more enjoyable.

[1083] The above is an embodiment of the present invention, and by viewing AR information provided through the smart glasses, a user can have a rich, informative and emotionally tailored video viewing experience.

[1084] The process flow will be explained below.

[1085] Step 1: The server acquires frames of a video, which allows the server to understand the scene of the video.

[1086] Step 2: The server processes the captured frames with image recognition algorithms, specifically analyzing features and patterns in the images to recognize buildings and the clothes worn by characters.

[1087] Step 3: The server retrieves related information from the database based on the recognition results, such as detailed information about the building, historical background, and the brand and styling of the clothing worn by the character.

[1088] Step 4: The server uses the emotion engine to recognize the user's emotions. Specifically, it analyzes emotions from the user's facial expressions and voice, and adjusts them to provide relevant information and AR displays that match the user's interests and preferences.

[1089] Step 5: The server combines the acquired related information with the emotion recognition results and processes the AR information. Specifically, the related information is packaged as AR data in the form of text and images, and the AR display is adjusted according to the user's emotions.

[1090] Step 6: The server sends the processed AR information to the user's smart glasses. Specifically, the server sends the AR data to the smart glasses using a communication protocol.

[1091] Step 7: The smart glasses (terminal) displays the received AR information. Specifically, by overlaying related information and AR display on the smart glasses display, the user can view the AR information while watching the video.

[1092] Step 8: The smart glasses recognize the user's emotions. Specifically, they analyze the user's emotions from their facial expressions and voice, and adjust the relevant information and AR display according to the user's emotions.

[1093] Example 2

[1094] Next, a description will be given of Example 2. In the following description, the data processing device 12 is referred to as a "server" and the robot 414 is referred to as a "terminal."

[1095] In conventional video viewing systems using visual devices, the means to provide information related to the video in real time are limited, limiting the user's viewing experience. In addition, there is a lack of technology to adjust related information taking into account the user's emotions. As a result, the information obtained while watching a video may not match the user's emotional state or interests.

[1096] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1097] In this invention, the server includes a means for performing image recognition of at least a structure or clothing worn by a character in a scene of a video, a means for acquiring related information from a data storage device based on a result of the image recognition, and a means for recognizing a user's emotion using an emotion engine and adjusting the related information and the augmented reality display in order to process the acquired related information in an augmented reality format. This makes it possible to provide related information in real time in augmented reality in accordance with the user's emotional state while watching a video.

[1098] A "visual device" is a digital device that enables a user to view images and videos, including, for example, smart glasses and head-mounted displays.

[1099] "Video" refers to visual content such as television, movies, video streaming services, etc., including content that a user views through a visual device.

[1100] "Related information" refers to knowledge or data related to a particular element in a video (e.g. a building, or the clothing worn by a character, etc.), including details, historical context, styling information, etc.

[1101] "Augmented reality" refers to a technology that overlays digital information on the user's field of vision, providing real-time information to the user's field of vision.

[1102] "Image recognition" refers to the use of computer vision techniques to identify and classify objects and features within a video scene.

[1103] "Data storage device" refers to a system that stores relevant information and makes it available for search and retrieval as needed, including databases and cloud storage.

[1104] An "emotion engine" refers to algorithms and technologies that analyze human facial expressions and behavior to infer their emotional state.

[1105] "Augmented reality processing" refers to the process of adjusting and transforming acquired data to fit the user's field of view, including adding visual effects and optimizing the way it is displayed.

[1106] "Real-time" refers to the fact that processing and display occur instantly at the moment the user begins viewing.

[1107] An embodiment of this system will now be described, which mainly comprises three parties: a server, a visual device (such as smart glasses), and a user.

[1108] Hardware and Software Configuration

[1109] server

[1110] The server acts as the main processing unit and uses the following hardware and software:

[1111] Hardware: A general server machine equipped with a high-performance CPU, GPU, and large memory capacity

[1112] software:

[1113] FFmpeg: A library for capturing video frames

[1114] TensorFlow: A library for implementing image recognition algorithms

[1115] Data storage device: Database management system such as MySQL or MongoDB

[1116] Emotion engine: Emotion recognition services such as Microsoft's Emotion API

[1117] AR development tools: Unity, Vuforia

[1118] Data processing and calculation

[1119] The server first captures video data frame by frame using FFmpeg. Next, it performs image recognition on these frames using TensorFlow. Based on the recognized information on buildings and clothes, it retrieves related information from databases such as MySQL and MongoDB. After that, it analyzes the user's emotions using Emotion API and processes the related information in an augmented reality format using Unity or Vuforia.

[1120] Smart glasses (terminal)

[1121] The smart glasses have the function of receiving augmented reality information sent from a server and displaying it over the user's field of vision. This is achieved by using the following hardware and software:

[1122] Hardware: Smart glasses with built-in camera, display sensor and processor

[1123] Software: Custom application for displaying augmented reality information

[1124] User Actions

[1125] The user wears the smart glasses and watches the video. Facial expression information captured by the smart glasses while watching the video is sent to the server in real time. The server analyzes this using an emotion engine, creates augmented reality information appropriate to the user's emotional state, and sends it back to the smart glasses. While watching the video, the user can visually obtain detailed information about buildings and clothes in real time.

[1126] Examples

[1127] For example, if a user is watching a historical movie, the server will recognize a building from the movie scene and retrieve the building's historical information from the database. At the same time, if the user's facial expression indicates that they are moved, the server will send augmented reality information with an emotional effect to the smart glasses. While watching the movie, the user can enjoy information such as the background and story of the building, making the experience richer.

[1128] Examples of prompt statements

[1129] Prompt: A user is watching a movie that contains a scene depicting a certain historical building. The user's emotion is recognized as being touched. Please suggest the most suitable AR information and its presentation effect that should be displayed in this situation.

[1130] The above is an embodiment of the present invention.

[1131] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1132] The flow of this system's program processing

[1133] Step 1: Grab a video frame

[1134] The server captures frames from the video the user is watching, using the FFmpeg library.

[1135] Input: The video the user is watching

[1136] Output: Captured frame image

[1137] Specific behavior:

[1138] The server runs the FFmpeg command to capture a frame every second: ffmpeg -i input.mp4 -vf "fps=1" frame_%04d.png.

[1139] Step 2: Applying image recognition algorithms

[1140] The server uses TensorFlow to perform image recognition on the captured frames, which allows it to recognize buildings and the clothes worn by characters.

[1141] Input: The captured frame image

[1142] Output: Recognized buildings and clothing information

[1143] Specific behavior:

[1144] The server loads a TensorFlow trained model and runs each frame through the model to obtain recognition results.

[1145] Example: recognition_result = tensorflow_model.predict(frame)

[1146] Step 3: Retrieving relevant information from the database

[1147] Based on the results of image recognition, the server retrieves relevant information about buildings and clothes from a database using MySQL and MongoDB.

[1148] Input: Recognized building and clothing information

[1149] Output: Related information (details, historical background, styling information)

[1150] Specific behavior:

[1151] The server retrieves the information by executing an SQL query like SELECT FROM building_info WHERE name = 'recognized_building'.

[1152] Step 4: Emotion recognition by emotion engine

[1153] The server uses an emotion engine such as Microsoft's Emotion API to recognize the user's emotions and receives the user's facial expression data from the smart glasses.

[1154] Input: User facial expression data

[1155] Output: The perceived emotional state of the user.

[1156] Specific behavior:

[1157] The server sends the user's facial expression data to the Emotion API and obtains the emotion analysis results.

[1158] Example: emotion_result = emotion_api.analyze(user_face_image)

[1159] Step 5: Processing the AR information

[1160] The server uses Unity or Vuforia to generate and process AR information based on the acquired related information and the user's emotional state.

[1161] Input: relevant information, the user's emotional state

[1162] Output: Processed AR information

[1163] Specific behavior:

[1164] The server generates 3D objects and effects in Unity and creates the AR display.

[1165] For users who are impressed, add specific effects (eg, fireworks or music).

[1166] Step 6: Sending information to the smart glasses

[1167] The server sends the processed AR information to the smart glasses, updating the information in real time using protocols such as WebSocket.

[1168] Input: Processed AR information

[1169] Output: Sending information to smart glasses

[1170] Specific behavior:

[1171] The server sends information using websocket.send(ar_information).

[1172] Step 7: Display and emotional feedback on smart glasses

[1173] The device (smart glasses) receives the AR information from the server and displays it over the user's field of view. At the same time, it recognizes the user's emotions again and updates the information as necessary.

[1174] Input: AR information sent from the server

[1175] Output: AR information displayed in the user's field of view, and the user's emotion data obtained again.

[1176] Specific behavior:

[1177] The smart glasses display the AR information received through a display application in the field of view.

[1178] The built-in camera is used to capture the user's facial expressions in real time and the data is sent to the server.

[1179] (Application example 2)

[1180] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal".

[1181] In conventional video viewing, there was a lack of a way to provide users with detailed information about the content in real time. In addition, there was no mechanism to adjust related information according to the user's emotions, so the viewing experience was uniform and it was not possible to provide information optimized for each individual user.

[1182] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means. In this invention, the server includes a means for performing image recognition of at least a building or clothes worn by a character in a scene of a video, a means for acquiring related information from a database based on the result of the image recognition, a means for processing the acquired related information in an AR format and transmitting it to the user's smart device, a means for recognizing the user's face using a camera of the smart device and analyzing emotions, and a means for adjusting the display content of the related information based on the analyzed emotions. This makes it possible to provide the user with an optimized viewing experience by providing the user with related information in real time and further adjusting the information according to the user's emotions.

[1183] A "smart device" is a device such as a smartphone, smart glasses, or a head-mounted display that is equipped with a camera to display images and recognize the user's face.

[1184] "Video" refers to dynamic video content, such as movies, television programs, and online streaming content.

[1185] "Building" refers to any structure or structure that appears in the video.

[1186] "Clothes worn by characters" refers to the clothing worn by people or characters appearing in the video.

[1187] An "image recognition means" is an algorithm or software that extracts information from frames of video and identifies specific objects or people.

[1188] "Related information" is supplemental data provided to users, such as detailed information about buildings or clothing, historical background, design information, etc.

[1189] A "database" is a collection of related information that can be stored and retrieved as needed.

[1190] "Means for processing in AR format" refers to a technology that uses augmented reality (AR) technology to superimpose related information on the user's field of vision.

[1191] A "transmitting means" is a communication means or protocol for sending data from the server to the smart device.

[1192] "Means for recognizing faces and analyzing emotions" refers to technology that captures the user's face with a camera and determines their emotional state through facial expression analysis, etc.

[1193] "Means for adjusting display content" refers to technology for changing the content or style of information displayed in AR based on recognized emotions.

[1194] The present invention is a system that provides related information in real time based on the content of a video when a user watches the video using a smart device, and further adjusts the information according to the user's emotions. The following describes how to implement the system in detail.

[1195] First, the server captures frames from the video and uses image recognition technology to identify buildings and character clothing. Image recognition is achieved using OpenCV, a commonly used image processing library, and an AI framework equipped with a deep learning model. The server also retrieves relevant information from a database and processes it into an AR format. The database pre-stores relevant information, but it is also possible to temporarily retrieve information via the Internet during processing.

[1196] Next, the smart device (e.g., smartphone, smart glasses) receives the AR information sent from the server and displays it over the user's field of view. This uses augmented reality (AR) technology such as Google ARCore. In addition, the camera of the smart device is used to recognize the user's face and perform emotion analysis. A pre-trained emotion recognition model (e.g., a deep learning model built using Keras) is used for emotion analysis. Once the emotion is analyzed, the display content of the AR information is adjusted based on the results. For example, if the user smiles, the color and display style of the related information are adjusted to be brighter.

[1197] As a concrete example, while a user is watching their favorite movie on their smartphone, behind-the-scenes and historical information about buildings and costumes that appear in the movie are displayed in real time in AR, and the information is adjusted according to the user's emotions: if the user smiles, the information display becomes more cheerful and bright.

[1198] The specific program processing for realizing this system is as follows.

[1199] 1. The server captures frames of video and uses image recognition technology to identify buildings and character clothing.

[1200] 2. Retrieve relevant information from the database, process it into AR format and send it to the smart device.

[1201] 3. Recognize the user's face using the smart device's camera and perform emotion analysis. The emotion recognition model used is built using Keras.

[1202] 4. The display content of the AR information is adjusted based on the emotion recognition results. For example, the color and style are adjusted to be brighter when the person is smiling.

[1203] 5. The smart device will then display this tailored information, providing the user with an improved viewing experience.

[1204] An example of a prompt for a generative AI model is as follows:

[1205] Suppose a user is watching a video on their smartphone. Information about buildings and costumes that appear in the movie scenes is retrieved in real time and displayed in AR format. Furthermore, the smartphone camera recognizes the user's emotions, and if the user smiles, the display of related information is adjusted to be more enjoyable.

[1206] The flow of the specific process in the application example 2 will be described with reference to FIG.

[1207] Step 1:

[1208] The server captures the video stream from the user's smart device in real-time. It extracts frames of the video and uses them for analysis. The input is the video stream and the output are the individual frames. This is done using an image processing library like OpenCV.

[1209] Step 2:

[1210] The server applies image recognition techniques to the extracted frames. Specifically, it uses a deep learning model to identify buildings and character clothing in the frames. The input is the captured frame, and the output is the identified building and clothing information. This is done using deep learning frameworks such as TensorFlow and PyTorch.

[1211] Step 3:

[1212] The server retrieves information related to the identified buildings or clothes from a database. The input is the identified object and the output is the related information. This can be done using a SQL or NoSQL database.

[1213] Step 4:

[1214] The server processes the retrieved relevant information into an augmented reality (AR) format. The input is the relevant information, and the output is the processed information for AR. This is achieved using an AR library such as Google ARCore.

[1215] Step 5:

[1216] The server sends the processed AR information to the user's smart device. The input is the processed AR information, and the output is the data sent to the user's smart device. This is done using REST APIs and WebSockets.

[1217] Step 6:

[1218] The device uses a camera to recognize the user's face. It then uses a facial recognition algorithm to analyze the user's emotions. The input is the camera image, and the output is the user's emotional information. Keras and OpenCV are used for emotion recognition.

[1219] Step 7:

[1220] The terminal adjusts the AR information received from the server based on the user's emotion information. For example, if the user is smiling, the color and style of the displayed information is adjusted to be brighter. The input is the user's emotion information and the processed AR information, and the output is the adjusted AR information.

[1221] Step 8:

[1222] The device displays the adjusted AR information overlaid on the user's field of view. The input is the adjusted AR information, and the output is the AR content displayed in the user's field of view. This is done using an AR framework such as ARCore or ARKit.

[1223] As a concrete example, when a user is watching a movie, detailed information about the buildings and clothes that appear in the scene is displayed, and if the user smiles, the information is adjusted to be more pleasant.

[1224] Suppose a user is watching a video on their smartphone. Information about buildings and costumes that appear in the movie scenes is retrieved in real time and displayed in AR format. Furthermore, the smartphone camera recognizes the user's emotions, and if the user smiles, the display of related information is adjusted to be more enjoyable.

[1225] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires a voice indicating a user input for the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1226] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by making a neural network perform deep learning. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating a voice, text data indicating a text, and image data indicating an image is input. The data generation model 58 performs inference on the input inference data according to the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1227] In the above embodiment, an example was given in which the specific process was performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the robot 414.

[1228] The emotion identification model 59 as an emotion engine may determine the emotion of the user according to a specific mapping. Specifically, the emotion identification model 59 may determine the emotion of the user according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the emotion of the robot, and the identification processing unit 290 may perform identification processing using the emotion of the robot.

[1229] FIG. 9 is a diagram showing an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive emotions are arranged. The more outside the concentric circles, the more emotions that represent states and actions that arise from a state of mind are arranged. Emotions are a concept that includes emotions and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions that occur in the brain are arranged. On the right side of the concentric circles, emotions that are generally induced by situational judgment are arranged. On the upper and lower sides of the concentric circles, emotions that are generally generated from reactions that occur in the brain and are induced by situational judgment are arranged. In addition, on the upper side of the concentric circles, emotions of "pleasure" are arranged, and on the lower side, emotions of "discomfort" are arranged. In this way, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1230] These emotions are distributed in the 3 o'clock direction of emotion map 400 and usually fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1231] The inside of emotion map 400 represents what is going on inside one's mind, and the outside of emotion map 400 represents behavior, so the further out you go on emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1232] Here, human emotions are based on various balances such as posture and blood sugar level, and when these balances are far from the ideal, it indicates an unpleasant state, and when they are close to the ideal, it indicates a pleasant state. Emotions can also be created for robots, cars, motorcycles, etc., based on various balances such as posture and battery level, so that when these balances are far from the ideal, it indicates an unpleasant state, and when they are close to the ideal, it indicates a pleasant state. The emotion map may be generated, for example, based on the emotion map of Dr. Mitsuyoshi (Research on speech emotion recognition and emotion brain physiological signal analysis system, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). On the left half of the emotion map, emotions belonging to an area called "reaction" where sensation is dominant are lined up. On the right half of the emotion map, emotions belonging to an area called "situation" where situation recognition is dominant are lined up.

[1233] The emotion map defines two emotions that promote learning. The first is the negative emotion around the middle of "repentance" or "remorse" on the situation side. In other words, this is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the positive emotion around "desire" on the response side. In other words, this is when the robot has positive feelings such as "I want more" or "I want to know more."

[1234] The emotion identification model 59 inputs the user input to a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the emotion of the user. This neural network is pre-trained based on multiple learning data that are combinations of the user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in Fig. 10. Fig. 10 shows an example in which multiple emotions, "relief," "calm," and "encouraging," have similar emotion values.

[1235] Although the system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, the system according to the present disclosure is not necessarily implemented in a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program that runs on a personal computer, or an application that runs on a smartphone or the like. The method according to the present disclosure may be provided to a user in the form of SaaS (Software as a Service).

[1236] In the above embodiment, an example is given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to input data.

[1237] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable computer-readable non-transitory storage medium such as a Universal Serial Bus (USB) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1238] In addition, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 upon request from the data processing device 12.

[1239] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1240] As the hardware resource for executing the specific process, various processors as shown below can be used. An example of the processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing the specific process by executing software, i.e., a program. Another example of the processor is a dedicated electric circuit, which is a processor having a circuit configuration designed exclusively for executing the specific process, such as a Field-Programmable Gate Array (FPGA), a Programmable Logic Device (PLD), or an Application Specific Integrated Circuit (ASIC). Each processor has a built-in or connected memory, and each processor executes the specific process by using the memory.

[1241] The hardware resource that executes the specific process may be one of these various processors, or may be a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific process may be a single processor.

[1242] As an example of a configuration using one processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a configuration using a processor that realizes the functions of the entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1243] Furthermore, more specifically, the hardware structure of these various processors can be an electric circuit that combines circuit elements such as semiconductor elements. The specific processes described above are merely examples. It goes without saying that unnecessary steps may be deleted, new steps may be added, or the order of processes may be changed without departing from the spirit of the invention.

[1244] The above description and illustrations are detailed descriptions of the parts related to the technology of the present disclosure, and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, function, action, and effect is an example of the configuration, function, action, and effect of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above description and illustrations, within the scope of the gist of the technology of the present disclosure. In addition, in order to avoid confusion and to facilitate understanding of the parts related to the technology of the present disclosure, the above description and illustrations omit explanations of technical common sense that do not require explanation in order to enable the implementation of the technology of the present disclosure.

[1245] All publications, patent applications, and standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or standard was specifically and individually indicated to be incorporated by reference.

[1246] The following is further disclosed regarding the above embodiment.

[1247] (Appendix 1) A system for providing AR-related information related to a video, including a television or movie, in real time through smart glasses when a user watches the video using the smart glasses, comprising: Performing image recognition of at least buildings or clothes worn by characters in scenes of the video; acquiring the related information from a database based on a result of the image recognition; The system processes the acquired related information in an AR format and transmits it to the user's smart glasses.

[1248] (Appendix 2) 2. The system of claim 1, wherein the system provides detailed information or historical context about buildings in scenes of the video in AR.

[1249] (Appendix 3) 2. The system of claim 1, wherein the system provides detailed information or styling information about clothing worn by a character in a scene of the video in AR.

[1250] (Appendix 4) 2. The system of claim 1, further comprising an emotion engine for recognizing emotions of the user.

[1251] (Appendix 5) 5. The system of claim 4, further comprising: a processor that recognizes an emotion of the user using the emotion engine; and adjusts the related information or the AR based on the recognition result.

[1252] (Appendix 6) 6. The system according to claim 5, wherein the emotion engine analyzes emotions from the user's facial expressions or voice, and provides the related information or the AR display in accordance with the user's interests or preferences.

[1253] "Example 1"

[1254] (Claim 1) A system for providing information related to a video being viewed by a user using a display device worn by the user in an augmented reality format, comprising: means for continuously acquiring frames of a video being viewed on said display device; means for applying an image recognition algorithm to the captured video frames to identify structures and clothing of people; A means for obtaining relevant information about the structure or the person's clothing from a database or a generative AI model based on the result of the image recognition; means for converting the acquired relevant information into an augmented reality format; means for transmitting the transformed augmented reality information to said display device; a means for displaying augmented reality information by overlaying it on a user's field of view using the display device; A system including:

[1255] (Claim 2) 13. The system of claim 1, providing detailed information or historical context about a building in a video scene in augmented reality.

[1256] (Claim 3) 10. The system of claim 1, providing detailed information or styling information about clothing of a person in a video scene in augmented reality.

[1257] "Application example 1"

[1258] (Claim 1) A system for providing, when a user is using smart glasses to perform a viewing experience, relevant information related to the viewing experience in real time through the smart glasses in AR, A means for image recognition of at least an object during the viewing experience; a means for acquiring the related information from a database or via the Internet based on a result of the image recognition; A means for processing the acquired related information in an AR format; means for transmitting the processed AR information to the smart glasses of the user; The smart glasses will display detailed product information and usage examples overlaid with AR, A system including:

[1259] (Claim 2) The system of claim 1, including an item recognition function and an image recognition algorithm.

[1260] (Claim 3) The system according to claim 1, comprising an information display function and an AR information generating means.

[1261] "Example 2 of combining emotion engines"

[1262] (Claim 1) A system for providing, when a user watches a video using a visual device, relevant information related to the video in real time through the visual device in augmented reality, comprising: A means for performing image recognition of at least a structure or clothing worn by a character in the scene of the video; means for acquiring the related information from a data storage device based on a result of the image recognition; a means for recognizing a user's emotion using an emotion engine and adjusting the related information and the augmented reality display, so as to process the acquired related information in an augmented reality format; means for transmitting the processed related information to the visual device of the user; a means for displaying the related information in the user's field of vision with the visual device and adjusting the display in accordance with the user's emotion; A system including:

[1263] (Claim 2) 10. The system of claim 1, further comprising: a display that displays augmented reality information about a structure in the video scene, the display displaying the structure being viewed;

[1264] (Claim 3) 10. The system of claim 1, providing detailed information or styling information in augmented reality about clothing worn by characters in the video scenes.

[1265] "Application example 2 when combining emotion engines"

[1266] (Claim 1) A system for providing related information related to a video including content in real time through a smart device when a user watches the video using the smart device, comprising: A means for performing image recognition of at least a building or clothes worn by a character in a scene of the video; means for acquiring the related information from a database based on a result of the image recognition; means for processing the acquired related information in an AR format and transmitting the processed related information to the smart device of the user; A means for recognizing a user's face and analyzing emotions using a camera of the smart device; a means for adjusting the display content of related information based on the analyzed emotion; A system including:

[1267] (Claim 2) The system of claim 1 , providing detailed information or historical background about buildings in the video scenes in an AR manner.

[1268] (Claim 3) The system of claim 1 , providing detailed information or styling information about clothes worn by a character in a scene of the video in AR. [Explanation of symbols]

[1269] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A system for providing information related to a video being viewed by a user using a display device worn by the user in an augmented reality format, comprising: A means for applying an image recognition algorithm to a video being viewed on the display device to perform image recognition and identifying at least a building or clothing of a person in a scene of the video; A means for obtaining relevant information about the structure or the person's clothing from a database or a generative AI model based on the result of the image recognition; means for converting the obtained related information into an augmented reality format; means for transmitting the converted augmented reality format information to the display device; means for displaying the augmented reality information on the display device in a manner superimposed on a user's field of view; A system including:

2. The system of claim 1 , further comprising: augmented reality provision of detailed information or historical context about buildings in video scenes.

3. The system of claim 1 , providing detailed information or styling information about clothing of a person in a video scene in augmented reality.

4. The system of claim 1 , further comprising: a user interface for recognizing an emotion of the user using an emotion engine for recognizing an emotion of the user; and adjusting the relevant information or the augmented reality information based on the recognition result.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A