system
The system addresses the challenge of managing and searching vast video data by automatically tagging and summarizing content, enabling efficient and personalized video retrieval.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-10
- Publication Date
- 2026-04-22
AI Technical Summary
The rapid increase in video data has made it difficult to efficiently organize and search video content, leading to challenges in quickly finding relevant information.
A system that analyzes video data to extract visual and textual features, automatically tags the data, and uses multiple modal learning technologies for high-speed and accurate searching, while generating video summaries to improve data management.
Enables efficient management and quick access to video content by automating the tagging process and providing personalized recommendations based on user queries and emotional states, reducing the burden of data management and improving user experience.
Smart Images

Figure 2026068316000001_ABST
Abstract
Description
Technical Field
[0001] The technology of this disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] With the rapid increase in video data, its management and retrieval have become difficult. Current technologies have the problem that it is difficult to efficiently organize video content and search based on content, and users cannot quickly find the information they need. By solving this problem, users need to be able to efficiently utilize a vast amount of video data.
Means for Solving the Problems
[0005] This invention provides a system that analyzes video data, extracts visual and textual features, and automatically tags them. This system utilizes multiple modal learning technologies to integrate the meaning of visual and textual information, enabling high-speed and accurate searching of related video data. This allows users to efficiently find the necessary video content based on specified queries. Furthermore, by generating video summaries using generative technologies, the system can improve the efficiency of data management by facilitating content comprehension.
[0006] "Video data" refers to digital data that includes visual information, and encompasses all types of video and video content.
[0007] "Means of receiving" refers to a function or device that acquires data from an external source and takes it internally in a format that can be processed.
[0008] "Means of analysis" refers to the techniques and methods used to break down and analyze received data and extract its content and characteristics.
[0009] "Visual features" refer to attributes that can be visually identified within an image or video, such as objects, shapes, and colors.
[0010] "Text information" refers to character information obtained from sources such as audio and subtitles, and is essentially digitized language data.
[0011] "Automatic tagging" refers to the process of automatically assigning relevant keywords and labels based on analyzed data.
[0012] A "database" is a computer system or software that systematically organizes, stores, and manages information, making it searchable and usable as needed.
[0013] A "search query" refers to a question or keyword that a user uses to obtain specific information.
[0014] "Related video data" refers to video data that is content-wise identical or correlated based on a search query or identified tags.
[0015] "User terminal" is a general term for electronic devices that a user operates to input and receive information.
[0016] "Modal learning" refers to a machine learning technique for integrating different information formats (e.g., vision and text) and analyzing meaning.
[0017] "Generation technology" refers to a technology for automatically generating new data and information using an AI model.
[0018] "Summary" refers to a form that extracts and concisely summarizes the important information of the original data.
Brief Description of Drawings
[0019] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9]Shows an emotion map to which a plurality of emotions are mapped. [Figure 10] Shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
Modes for Carrying Out the Invention
[0020] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0021] First, the language used in the following description will be described.
[0022] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be one arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be one type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0023] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0024] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0025] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0027] [First Embodiment]
[0028] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0029] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0030] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0031] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0032] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0034] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0035] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0036] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0037] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0038] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0039] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0040] Embodiments of the present invention are systems comprising a user-operated terminal, a central processing server, and a network that connects these. This system automatically analyzes and tags video data provided by the user, streamlining subsequent searching and management.
[0041] When a user uploads video data to the system via their device, the server receives the data. The server uses an AI model to analyze the visual features of the video and extract related text information. This analysis automatically tags the video, and the information is stored in a database.
[0042] This tagged data allows users to enter search queries from their devices and quickly find corresponding video content. The search is performed by the server associating the tag information of relevant videos with the search query, providing the user with the most suitable results.
[0043] For example, if a user uploads video data to the system that includes "a family's daily life filmed at the zoo," the server automatically generates tags such as "animals," "family," "children," and "outdoors" from this video and registers them in the database. Later, when the user performs a search using the query "children and animals," the server quickly searches for relevant videos and provides the results to the terminal.
[0044] Furthermore, this system uses generation technology to create summaries of selected videos, helping users effectively manage content. These summaries include important scenes and key phrases from the video, allowing users to instantly grasp the content.
[0045] These features allow users to efficiently manage vast amounts of video data and quickly access the information they need. Furthermore, this system reduces the burden of data management, contributing to improved efficiency in work and daily life.
[0046] The following describes the processing flow.
[0047] Step 1:
[0048] The user uses their device to select video data and perform the upload operation. At that time, preparations are made to send the data to the server via the internet.
[0049] Step 2:
[0050] The device sends the video data to be uploaded to the server. Once the data transmission is complete, the user is notified that the transmission is complete.
[0051] Step 3:
[0052] The server then processes the received video data. First, the data is input into an AI model, and visual features are extracted using image recognition technology.
[0053] Step 4:
[0054] The server uses speech recognition technology to extract text information from the audio of the video. This text information is useful for data retrieval in later processes.
[0055] Step 5:
[0056] The server integrates extracted visual features and text information using multiple modal learning technologies to gain a deeper understanding of the meaning of the video content.
[0057] Step 6:
[0058] The server automatically generates relevant tags based on the information obtained through analysis. This tagging will improve future search performance.
[0059] Step 7:
[0060] The server saves the tagged video data to a database, making it searchable. Once saving is complete, the results are recorded as a log.
[0061] Step 8:
[0062] The user enters a search query through their device and requests to search for a specific video. The device then sends this search request to the server.
[0063] Step 9:
[0064] The server analyzes the received search query and searches the database for video data related to the query. During the matching process, it understands the meaning of the search query and selects the most suitable search results based on tag information.
[0065] Step 10:
[0066] The server lists the search results and sends them back to the terminal. The user then reviews the results visually in an easy-to-understand format on the terminal.
[0067] Step 11:
[0068] If a user requires more detailed information, they can request a summary of a specific video through their device.
[0069] Step 12:
[0070] The server uses generation technology to automatically generate a summary of the specified video content and sends that content to the terminal.
[0071] Step 13:
[0072] Users can view the summary generated on their device and quickly understand the content.
[0073] These steps enable users to efficiently search and manage video data and quickly access the information they need.
[0074] (Example 1)
[0075] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0076] There is a need to efficiently analyze vast amounts of video data and quickly search for and provide relevant information. However, conventional methods require significant time and effort for manual tagging and searching, making efficient data management and information access difficult. Furthermore, extracting necessary information requires advanced analysis and summarization, which needs to be automated.
[0077] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0078] In this invention, the server includes a device for receiving video data, a device for analyzing the received video data and extracting visual attributes and text information, and a device for automatically tagging based on those attributes. This automates the analysis of vast amounts of video data, enabling efficient tagging and searching.
[0079] "Video data" refers to a series of images or video data containing visual information, which are recorded or transmitted electronically.
[0080] "Device" refers to a system or part of a hardware or software designed to perform a specific function or process.
[0081] "Analysis" refers to the process of examining data or information in detail to understand its structure and meaning.
[0082] "Visual attributes" refer to visually-based features such as objects, shapes, and colors present within video data.
[0083] "Text information" refers to information consisting of characters and words related to video data, including descriptions and tags extracted from visual elements.
[0084] "Automatic tagging" refers to the process of mechanically adding keywords and categories related to video data using AI or algorithms.
[0085] "Information recording media" refers to physical or virtual storage used to store and maintain digital data over long periods of time.
[0086] A "search request" refers to a query or question entered by a user to retrieve specific information.
[0087] A "user-operated device" refers to a computer or portable device that a human user directly operates and uses to interact with a system.
[0088] "Generative technologies" refer to a set of technical methods or tools for generating automatically generated data and information.
[0089] Embodiments of this invention include a system comprising a user-operated terminal, a central data processing server, and a network connecting them. The aim of this system is to efficiently manage vast amounts of video data and enable rapid searching and access to information.
[0090] The user operates a terminal to upload video data to the system. The terminal sends the data to the server via the internet. The server inputs the video data into an AI model using Python (e.g., TENSORFLOW® or PyTorch) and performs the extraction of visual attributes and text information. This AI model performs image analysis on each video frame, generating tags by detecting objects and recognizing scenes.
[0091] The server automatically tags the video data using the generated tags and stores it in an information storage medium, such as a NoSQL database (e.g., MongoDB). This stored data can be quickly accessed in response to subsequent search requests.
[0092] The user sends a search request using keywords from their device to the server. The server executes the search request against the database and generates a list of video data that matches the query. This result is sent back to the user's device, and the user can view the results directly.
[0093] Furthermore, the server uses generation technology to generate a summary of the selected video data. This summary utilizes a natural language processing model (e.g., using Hugging Face's Transformers) to transcribe important scenes and key phrases from the video into text, allowing the user to quickly grasp the content.
[0094] For example, if a user uploads "a video of their family taken at a zoo," the server generates tags such as "animals," "family," "children," and "outdoors." This tag information is stored in a database and used later when a search request for "children and animals" is received. The server quickly searches for relevant videos and provides the results to the user's device.
[0095] An example of a prompt message is, "Analyze the following video and generate relevant tags. Based on the features in the video, prioritize extracting tags related to family or animals." This prompt is used as input to the AI model, which helps in detailed tag generation and efficient data management.
[0096] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0097] Step 1:
[0098] The user uses their device to select and upload the video data they want to analyze. The device then takes the video data as input and sends it to the server via the network. During this process, a progress bar is displayed to show the status of data transmission, giving the user peace of mind.
[0099] Step 2:
[0100] The server receives video data transmitted from the terminal and performs a hash check to ensure data integrity. The received video data becomes the input for the analysis process. The server feeds the video data into an AI model (e.g., using TensorFlow or PyTorch) and extracts visual features for each image frame. This process performs object detection and scene recognition, and outputs extracted visual attributes and text information.
[0101] Step 3:
[0102] The server automatically tags the video based on visual features extracted by the AI model. The generated tags indicate the content of the video, such as "animals" or "children." These tags are added to the video data, resulting in tagged data output. The server stores this tagged data on an information storage medium. A NoSQL database (e.g., MongoDB) is used for storage to ensure fast access and scalability.
[0103] Step 4:
[0104] The user enters a keyword-based search request from their device and sends it to the server. The search query arrives at the server as input. The server parses the query and matches it with tag information in the database. This process extracts matching video data and outputs a list of search results. This list is then sent back to the user's device.
[0105] Step 5:
[0106] The server uses generative techniques to create a summary from the video data selected by the user. A natural language processing model (e.g., using Hugging Face's Transformers) extracts important scenes and key phrases from the video and provides output as a text summary. The server sends this summary information to the user's terminal, allowing the user to easily understand the content of the video.
[0107] (Application Example 1)
[0108] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0109] In today's world, with the ever-increasing volume of video content, users are required to quickly search for and view the information they desire from a vast amount of video data. However, manual tagging and classification are labor-intensive, making efficient management difficult. Furthermore, the lack of personalized recommendation features means users may miss out on content that interests them. It is necessary to solve these problems and provide an environment where users can enjoy video content efficiently and accurately.
[0110] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0111] In this invention, the server includes a device for extracting visual features and symbolic information, a device for automatically tagging content, and a device for providing personalized recommendations using viewing history. This enables users to efficiently manage vast amounts of video content and receive personalized recommendations based on visual features and viewing history.
[0112] "Video data" refers to digital information that records visual information continuously over time.
[0113] A "device" is a set of hardware or software designed to perform a specific function.
[0114] "Analysis" is the process of breaking down data and information into detailed components and understanding and evaluating their elements and structure.
[0115] "Visual features" are attributes that are identified by their appearance, such as color, shape, and pattern, contained within an image or video.
[0116] "Symbolic information" refers to text data that extracts specific meanings and content from video data and explains them.
[0117] "Tagging" is the process of assigning identifying labels or keywords to data.
[0118] A "storage device" is a medium or device capable of storing information and data for a long period of time.
[0119] "Search input" refers to the provision of queries or keywords that a user makes to the system in order to find specific data.
[0120] An "information processing system" is a set of computer systems configured to collect, store, analyze, and process data in order to provide necessary information.
[0121] A "recommendation list" is a list that presents highly relevant options based on the user's interests and past behavior.
[0122] "Personalized recommendations" are a process that presents the most suitable options based on the user's specific interests and preferences.
[0123] The system for implementing this invention is designed to efficiently manage video data and provide that data to users in a useful format. The method for implementing this system is described below.
[0124] The server first analyzes the video data received from the user's device. This analysis uses deep learning libraries such as TensorFlow to extract visual features and simultaneously generate symbolic information. Based on this information, the system automatically tags the data and stores it in storage using MongoDB.
[0125] Next, the server receives a search input from the user. This search query can be received in the form of a prompt message, for example, "I'm looking for recommended videos that contain both comedy and heartwarming scenes." The server then uses ElasticSearch® to quickly search for relevant video data and generate the best possible search results.
[0126] Furthermore, the server considers viewing history and generates personalized recommendation lists. These lists are displayed by an application built with React Native on the user's device, making it easy for users to find content that matches their interests.
[0127] For example, when a user searches for a comedy movie on the weekend, the system adds relevant movies to a recommendation list based on the user's past viewing and search history, and also presents videos that other viewers have given high ratings to. This allows users to select the content that best suits them from a vast amount of video data.
[0128] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0129] Step 1:
[0130] The user uploads video data to the system using their device. The uploaded video data is sent to the server. This data becomes the input for subsequent processing.
[0131] Step 2:
[0132] The server analyzes the received video data using TensorFlow. The analysis extracts visual features from the video and generates related symbolic information (text) based on these features. This process provides specific tag information for the video data.
[0133] Step 3:
[0134] The server automatically tags the extracted visual features and symbolic information. These tags are stored in MongoDB and output as a tagged database that will serve as the basis for future searches and recommendations.
[0135] Step 4:
[0136] The user enters a search query from their device. The query is received as a prompt that can be processed using natural language processing technology (e.g., "I'm looking for recommended videos that contain both comedy and heartwarming scenes."). This prompt becomes the input to the server.
[0137] Step 5:
[0138] The server uses Elasticsearch to match tag information in the database with search queries and quickly retrieve relevant video data. This process outputs search results for the video that best suits the user's request.
[0139] Step 6:
[0140] The server generates personalized recommendation lists based on viewing history and search queries. To achieve this, it analyzes past viewing activity and selects highly relevant content. This process results in an optimized video recommendation list for the user.
[0141] Step 7:
[0142] The generated search results and recommendation lists are displayed in a device application built with React Native. Users can then select and view video content that interests them based on the information presented.
[0143] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0144] Embodiments of the present invention include a system incorporating a user terminal, a central processing server, a network connecting them, and an emotion engine that recognizes the user's emotional state. This system allows the user to efficiently search for and manage video data corresponding to their emotions.
[0145] When a user selects video data through their device and uploads it to the system, this data is sent to a server. The server analyzes the received data using an AI model and extracts visual features and text information. Based on this information, automatic tagging is performed and stored in a database. This process enables efficient searching of video content.
[0146] Furthermore, the system incorporates an emotion engine that analyzes the user's emotions in real time while they are viewing videos. The emotion engine reads the user's emotional state from their facial expressions and voice and sends this information to the server. The server can then use this emotion information to recommend video content that is appropriate for the user's emotions and to adjust the search results.
[0147] As a concrete example, when a user views "soothing animal videos," the system's emotion engine analyzes the user's facial expressions and reactions. If the user appears relaxed and happy, the server prioritizes displaying similar relaxing video content. This process allows users to easily find video content that suits their emotional state.
[0148] Furthermore, by incorporating analysis information from the emotion engine into the tagging process, the tagging accuracy of video data is improved, resulting in more personalized search results. This makes it possible to provide users with the information they are looking for more quickly and accurately.
[0149] In this way, this system enables flexible management and retrieval of video data that takes into account the user's emotional state, thereby improving the user experience. By having the entire system work together to provide information that meets the user's needs, it can significantly improve efficiency in work and personal activities.
[0150] The following describes the processing flow.
[0151] Step 1:
[0152] The user uses a terminal to select video data and initiate the upload to the system. The terminal prepares the video data and transfers it to the server via the network.
[0153] Step 2:
[0154] The server processes the video data received from the terminal, using an AI model to extract visual features from the data. These extracted features are then used for tagging.
[0155] Step 3:
[0156] The server converts the audio information from the video data into text and analyzes the content within the video. This analysis is performed to generate highly relevant tags.
[0157] Step 4:
[0158] The server integrates visual features and text information and performs automatic tagging. This process assigns keywords related to the video data and stores them in a database.
[0159] Step 5:
[0160] When a user operates the device and begins watching a video, the emotion engine activates and starts analyzing the user's facial expressions and vocal characteristics.
[0161] Step 6:
[0162] The emotion engine analyzes the user's emotional state in real time and sends the obtained emotional information to the server.
[0163] Step 7:
[0164] The server selects video content that matches the user's emotions based on the received emotional information. In this process, videos with tags that match the user's emotions are given priority.
[0165] Step 8:
[0166] The user enters a search query on their device to search for a specific video. The device then sends this query to the server.
[0167] Step 9:
[0168] The server integrates search queries, tagging information, and user sentiment information to retrieve the most suitable video data from the database.
[0169] Step 10:
[0170] The search results are listed, ranked according to the user's emotional state, and sent to the device.
[0171] Step 11:
[0172] The terminal displays ranked search results received from the server to the user, presenting them in a visually easy-to-understand format.
[0173] Step 12:
[0174] If a user needs a summary of a specific video, they can request the server to generate the summary through their device.
[0175] Step 13:
[0176] The server uses generation technology to create a summary of the selected video content and provides it to the terminal.
[0177] Step 14:
[0178] Users can view summaries generated on their devices and quickly grasp the necessary information.
[0179] These steps enable users to efficiently search for and watch videos that resonate with their emotions, and to quickly understand the content they need.
[0180] (Example 2)
[0181] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0182] In today's digital society, it is crucial for users to quickly and accurately find appropriate digital content that matches their individual emotional state. However, conventional systems lacked the means to analyze user emotions in real time and recommend appropriate content, hindering improvements in the user experience. Furthermore, systems capable of dynamic content recommendation based on emotions were limited, and there was a need for information provision tailored to the individual needs of users.
[0183] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0184] In this invention, the server includes means for providing a terminal for users to upload data, means for using machine learning techniques to analyze the received data and extract its characteristics, means for automatically associating and storing identification information with the data, means for analyzing the user's emotional state, means for selecting and providing relevant information based on the emotional state, and means for dynamically adjusting the search results of the data according to the user's emotions. This enables advanced data recommendation and search that takes into account the user's emotional state, resulting in a significant improvement in the user experience.
[0185] "A terminal for users to upload data" refers to an electronic device used to transmit digital data specified by the user to the system.
[0186] "Means of using machine learning technology" refers to technologies that apply artificial intelligence algorithms to analyze received digital data and automatically identify and extract visual and textual information.
[0187] "Means for automatically associating and storing identification information with data" refers to the process of automatically assigning tags and metadata to analyzed digital data and saving that data to a storage device such as a database.
[0188] "Means for analyzing a user's emotional state" refers to technologies that analyze data such as a user's facial expressions and voice to identify the emotions they are expressing.
[0189] "Means of selecting and providing relevant information based on the aforementioned emotional state" refers to a process that automatically selects and presents digital content appropriate to the user's emotions based on the analyzed emotions of that user.
[0190] "Means of dynamically adjusting search results in response to user sentiment" refers to a mechanism that takes user sentiment data into consideration, modifies normal search results in real time, and prioritizes displaying the information most relevant to the user.
[0191] To implement this invention, a system is required in which a user, a terminal, and a server work together. The user can access the system using the terminal and select video data that they wish to view or upload. The terminal has an intuitive interface, making it easy for the user to operate.
[0192] In analyzing video data, the server utilizes machine learning techniques to analyze the received video data. Specifically, it can use widely used deep learning frameworks such as TensorFlow or PyTorch. The server uses these to extract visual features and text information from the video and automatically assigns identification information as tags. This automated tagging process allows the data to be efficiently stored in a database, enabling rapid searching and retrieval later on.
[0193] Furthermore, the server is equipped with an emotion engine to analyze the user's emotional state. This engine analyzes facial expressions and voice data transmitted from the user's device, allowing it to understand the user's emotional state in real time. Facial recognition technology and voice analysis technology are used for the analysis.
[0194] Based on the acquired emotional information, the server recommends video content that is appropriate for the user's current mood. In this recommendation process, the content recommendation algorithm reflects the emotional data and prioritizes providing the most relevant information to the user. For example, if a user is relaxing after watching "soothing animal videos," the server will continue to prioritize displaying videos with similar relaxing effects.
[0195] As a concrete example, here is an example of a prompt statement that utilizes a generative AI model:
[0196] "Please search for relaxing animal videos. If the user smiles, please recommend similar videos."
[0197] This significantly improves the user experience and enables content viewing tailored to individual emotional states. The integrated operation of the entire system makes managing and searching for digital content even more efficient.
[0198] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0199] Step 1:
[0200] The user selects digital data and uploads it to the system from the terminal. The input is a video file selected by the user, which is sent to the server via the terminal's interface. The output is the data arriving at the server. The terminal provides the user with a simple and intuitive UI, making file selection and uploading easy.
[0201] Step 2:
[0202] The server analyzes the received video data. The input is a video file, which is then fed into an AI model to extract visual features and textual information. Deep learning frameworks such as TensorFlow and PyTorch are used for data processing, and automatically tagged data is generated as output. The server efficiently stores this information in a database.
[0203] Step 3:
[0204] The system receives data to analyze the user's emotional state in real time. Input includes facial expressions and voice data sent from the device, which the emotion engine processes to determine the user's emotional state. Emotional information is generated as output and sent to the server. Facial recognition and voice tone analysis technologies are applied in the analysis.
[0205] Step 4:
[0206] The server uses user sentiment information to recommend content. It uses sentiment information and video content from its database as input. A content recommendation algorithm selects videos that match the sentiment, generating a list of candidates to display to the user. The server then sends this list to the user's device.
[0207] Step 5:
[0208] The user views recommended video content on their device. The input is a list of content sent from the server, from which the user can select and begin playback. The output further collects the user's emotional state and sends it back to the system as feedback to help improve future recommendations. This allows for continuous improvement of the user experience.
[0209] (Application Example 2)
[0210] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0211] Modern video streaming services are required to appropriately select and efficiently deliver the video content that users want, but content recommendations that take into account the user's emotional state are not being adequately implemented. As a result, user satisfaction is low, and the value of video streaming services is not being fully realized.
[0212] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0213] In this invention, the server includes means for analyzing the user's facial expression and voice information to determine their emotional state, means for recommending video content based on the determined emotional state, and means for providing search results to the user's terminal. This makes it possible to recommend video content based on the user's emotional state, enabling the efficient and effective provision of the video content the user desires.
[0214] "Video data" refers to digital data containing visual information, such as videos and images.
[0215] "Means of receiving" refers to mechanisms and methods for taking in data from an external source, such as acquiring information via networks or communication devices.
[0216] "Methods for analysis and extraction of visual and textual features" refers to methods for analyzing the content of video data and extracting features from image and textual information.
[0217] "Methods for automatic tagging" refer to the process of automatically assigning identifiable words or labels to data based on extracted features.
[0218] "Means of storing information in a database" refers to a system for long-term storage of structured information.
[0219] "A means of receiving search queries and searching for relevant video data" refers to a method of taking in information requests from users and finding corresponding video data.
[0220] "Means of providing to the user's terminal" refers to functions for displaying or delivering search results and recommended content to the user's device.
[0221] "A means of analyzing a user's facial expression and voice information to determine their emotional state" refers to a process of identifying a user's emotions at a given time by analyzing their facial movements and tone of voice.
[0222] "Methods for recommending video content based on emotional state" refer to technologies that take into account the determined emotions and select and present appropriate videos.
[0223] This system is implemented through a network that includes the user's smart device and a server. Video data viewed by the user is transmitted to the server in real time via a device such as a smartphone. The server analyzes the received video data and extracts visual features and text information. Feature extraction is performed using deep learning libraries such as OpenCV and TensorFlow.
[0224] Based on the analyzed features, a generative AI model automatically tags the data and stores it in a structured database. Furthermore, software equipped with an emotion engine detects changes in the user's facial expressions and voice to determine their emotional state. This analysis utilizes the smartphone's hardware, including its camera and microphone.
[0225] Once the user's emotional state is determined, the system uses that information to recommend video content that best suits the user's current emotions. This makes it easy for users to find content that resonates with their feelings.
[0226] For example, if the system detects that a user is crying while watching an emotionally moving film, it will recommend another film that is both touching and heartwarming. If the system detects laughter in the user's voice, it will then display a more entertaining comedy.
[0227] Here are some examples of prompts generated using a generative AI model. Instructions such as "If the user smiles, please display the following recommended content" or "If the user shows an emotion such as crying, please recommend a heartwarming movie" are used.
[0228] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0229] Step 1:
[0230] The user selects the video content they want to watch via their smart device. The device receives the video data and sends it to the server in real time. In this step, the video file is the input, and the output is data sent to the server in streaming format.
[0231] Step 2:
[0232] The server uses an AI model to extract visual features and text information to analyze the received video data. The input is video data, and the output is a list of features. Image analysis and natural language processing are performed using libraries such as OpenCV and TensorFlow, and the features are extracted as basic information for automatic tagging.
[0233] Step 3:
[0234] The server automatically tags the extracted visual and text features using a generative AI model. This tagging information is stored in a database. The input is a list of features, and the output is a data structure with automatic tags. The database storage process takes place at this stage.
[0235] Step 4:
[0236] When a user views video using their device, the device uses its camera and microphone to collect facial and audio information in real time. The collected data is sent to a server. The input consists of facial and audio data, and the output is the transmission of a data stream to the server.
[0237] Step 5:
[0238] The server analyzes the received facial expression and voice information to determine the user's emotional state. The input consists of facial expression data and voice data, and the output is a determination result indicating the emotional state (e.g., joy, sadness, etc.). The emotion engine performs this data calculation.
[0239] Step 6:
[0240] The server recommends video content that matches the user's current mood based on the determined emotional state. The input is the emotional state determination result and video database information, and the output is a list of recommended video content. A generative AI model is used for this recommendation process.
[0241] Step 7:
[0242] The terminal provides the user with video content recommended by the server. The user selects the content they want to watch next from the recommendation list displayed on the terminal. The input is the recommendation list received from the server, and the output is the video content selected by the user.
[0243] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0244] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0245] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0246] [Second Embodiment]
[0247] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0248] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0249] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0250] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0251] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0252] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0253] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0254] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0255] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0256] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0257] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0258] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0259] Embodiments of the present invention are systems comprising a user-operated terminal, a central processing server, and a network that connects these. This system automatically analyzes and tags video data provided by the user, streamlining subsequent searching and management.
[0260] When a user uploads video data to the system via their device, the server receives the data. The server uses an AI model to analyze the visual features of the video and extract related text information. This analysis automatically tags the video, and the information is stored in a database.
[0261] This tagged data allows users to enter search queries from their devices and quickly find corresponding video content. The search is performed by the server associating the tag information of relevant videos with the search query, providing the user with the most suitable results.
[0262] For example, if a user uploads video data to the system that includes "a family's daily life filmed at the zoo," the server automatically generates tags such as "animals," "family," "children," and "outdoors" from this video and registers them in the database. Later, when the user performs a search using the query "children and animals," the server quickly searches for relevant videos and provides the results to the terminal.
[0263] Furthermore, this system uses generation technology to create summaries of selected videos, helping users effectively manage content. These summaries include important scenes and key phrases from the video, allowing users to instantly grasp the content.
[0264] These features allow users to efficiently manage vast amounts of video data and quickly access the information they need. Furthermore, this system reduces the burden of data management, contributing to improved efficiency in work and daily life.
[0265] The following describes the processing flow.
[0266] Step 1:
[0267] The user uses their device to select video data and perform the upload operation. At that time, preparations are made to send the data to the server via the internet.
[0268] Step 2:
[0269] The device sends the video data to be uploaded to the server. Once the data transmission is complete, the user is notified that the transmission is complete.
[0270] Step 3:
[0271] The server then processes the received video data. First, the data is input into an AI model, and visual features are extracted using image recognition technology.
[0272] Step 4:
[0273] The server uses speech recognition technology to extract text information from the audio of the video. This text information is useful for data retrieval in later processes.
[0274] Step 5:
[0275] The server integrates extracted visual features and text information using multiple modal learning technologies to gain a deeper understanding of the meaning of the video content.
[0276] Step 6:
[0277] The server automatically generates relevant tags based on the information obtained through analysis. This tagging will improve future search performance.
[0278] Step 7:
[0279] The server saves the tagged video data to a database, making it searchable. Once saving is complete, the results are recorded as a log.
[0280] Step 8:
[0281] The user enters a search query through their device and requests to search for a specific video. The device then sends this search request to the server.
[0282] Step 9:
[0283] The server analyzes the received search query and searches the database for video data related to the query. When making an association, it understands the meaning of the search query and selects the optimal search results based on the tag information.
[0284] Step 10:
[0285] The server lists the search results and sends them back to the terminal. The user can view the results in a visually understandable form through the terminal.
[0286] Step 11:
[0287] If the user needs more detailed information, they request a summary of a specific video through the terminal.
[0288] Step 12:
[0289] The server automatically generates a summary of the specified video content using generation technology and sends the content to the terminal.
[0290] Step 13:
[0291] The user can view the summary generated on the terminal and quickly understand the content of the video.
[0292] Through these steps, the user can efficiently search for and manage video data and quickly access the necessary information.
[0293] (Example 1)
[0294] Next, Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal". <000093There is a need to efficiently analyze vast amounts of video data and quickly search for and provide relevant information. However, conventional methods require significant time and effort for manual tagging and searching, making efficient data management and information access difficult. Furthermore, extracting necessary information requires advanced analysis and summarization, which needs to be automated.
[0296] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0297] In this invention, the server includes a device for receiving video data, a device for analyzing the received video data and extracting visual attributes and text information, and a device for automatically tagging based on those attributes. This automates the analysis of vast amounts of video data, enabling efficient tagging and searching.
[0298] "Video data" refers to a series of images or video data containing visual information, which are recorded or transmitted electronically.
[0299] "Device" refers to a system or part of a hardware or software designed to perform a specific function or process.
[0300] "Analysis" refers to the process of examining data or information in detail to understand its structure and meaning.
[0301] "Visual attributes" refer to visually-based features such as objects, shapes, and colors present within video data.
[0302] "Text information" refers to information consisting of characters and words related to video data, including descriptions and tags extracted from visual elements.
[0303] "Automatic tagging" refers to the process of mechanically adding keywords and categories related to video data using AI or algorithms.
[0304] "Information recording medium" refers to physical or virtual storage used to store and maintain digital data for a long time.
[0305] "Search request" refers to a query or question entered by a user to obtain specific information.
[0306] "User operation device" refers to a computer or mobile device directly operated by a human user to interact with the system.
[0307] "Generation technology" refers to a group of technical methods or tools for generating automatically generated data and information.
[0308] An embodiment of this invention is a system including a terminal operated by a user, a server that performs data processing centrally, and a network connecting them. The purpose of this system is to efficiently manage a huge amount of video data and enable quick search and access to information.
[0309] The user operates the terminal to upload video data to the system. The terminal transmits the data to the server via the Internet. The server inputs the video data into an AI model (e.g., TensorFlow or PyTorch) using Python and executes extraction of visual attributes and text information. This AI model performs image analysis for each video frame and generates tags by performing object detection and scene recognition.
[0310] The server automatically tags the video data using the generated tags and stores it in an information recording medium, such as a NoSQL database (e.g., MongoDB). This stored data can be quickly accessed in response to a later search request.
[0311] The user sends a search request using keywords from their device to the server. The server executes the search request against the database and generates a list of video data that matches the query. This result is sent back to the user's device, and the user can view the results directly.
[0312] Furthermore, the server uses generation technology to generate a summary of the selected video data. This summary utilizes a natural language processing model (e.g., using Hugging Face's Transformers) to transcribe important scenes and key phrases from the video into text, allowing the user to quickly grasp the content.
[0313] For example, if a user uploads "a video of their family taken at a zoo," the server generates tags such as "animals," "family," "children," and "outdoors." This tag information is stored in a database and used later when a search request for "children and animals" is received. The server quickly searches for relevant videos and provides the results to the user's device.
[0314] An example of a prompt message is, "Analyze the following video and generate relevant tags. Based on the features in the video, prioritize extracting tags related to family or animals." This prompt is used as input to the AI model, which helps in detailed tag generation and efficient data management.
[0315] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0316] Step 1:
[0317] The user uses their device to select and upload the video data they want to analyze. The device then takes the video data as input and sends it to the server via the network. During this process, a progress bar is displayed to show the status of data transmission, giving the user peace of mind.
[0318] Step 2:
[0319] The server receives video data transmitted from the terminal and performs a hash check to ensure data integrity. The received video data becomes the input for the analysis process. The server feeds the video data into an AI model (e.g., using TensorFlow or PyTorch) and extracts visual features for each image frame. This process performs object detection and scene recognition, and outputs extracted visual attributes and text information.
[0320] Step 3:
[0321] The server automatically tags the video based on visual features extracted by the AI model. The generated tags indicate the content of the video, such as "animals" or "children." These tags are added to the video data, resulting in tagged data output. The server stores this tagged data on an information storage medium. A NoSQL database (e.g., MongoDB) is used for storage to ensure fast access and scalability.
[0322] Step 4:
[0323] The user enters a keyword-based search request from their device and sends it to the server. The search query arrives at the server as input. The server parses the query and matches it with tag information in the database. This process extracts matching video data and outputs a list of search results. This list is then sent back to the user's device.
[0324] Step 5:
[0325] The server uses generative techniques to create a summary from the video data selected by the user. A natural language processing model (e.g., using Hugging Face's Transformers) extracts important scenes and key phrases from the video and provides output as a text summary. The server sends this summary information to the user's terminal, allowing the user to easily understand the content of the video.
[0326] (Application Example 1)
[0327] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0328] In today's world, with the ever-increasing volume of video content, users are required to quickly search for and view the information they desire from a vast amount of video data. However, manual tagging and classification are labor-intensive, making efficient management difficult. Furthermore, the lack of personalized recommendation features means users may miss out on content that interests them. It is necessary to solve these problems and provide an environment where users can enjoy video content efficiently and accurately.
[0329] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0330] In this invention, the server includes a device for extracting visual features and symbolic information, a device for automatically tagging content, and a device for providing personalized recommendations using viewing history. This enables users to efficiently manage vast amounts of video content and receive personalized recommendations based on visual features and viewing history.
[0331] "Video data" refers to digital information that records visual information continuously over time.
[0332] A "device" is a set of hardware or software designed to perform a specific function.
[0333] "Analysis" is the process of breaking down data and information into detailed components and understanding and evaluating their elements and structure.
[0334] "Visual features" are attributes that are identified by their appearance, such as color, shape, and pattern, contained within an image or video.
[0335] "Symbolic information" refers to text data that extracts specific meanings and content from video data and explains them.
[0336] "Tagging" is the process of assigning identifying labels or keywords to data.
[0337] A "storage device" is a medium or device capable of storing information and data for a long period of time.
[0338] "Search input" refers to the provision of queries or keywords that a user makes to the system in order to find specific data.
[0339] An "information processing system" is a set of computer systems configured to collect, store, analyze, and process data in order to provide necessary information.
[0340] A "recommendation list" is a list that presents highly relevant options based on the user's interests and past behavior.
[0341] "Personalized recommendations" are a process that presents the most suitable options based on the user's specific interests and preferences.
[0342] The system for implementing this invention is designed to efficiently manage video data and provide that data to users in a useful format. The method for implementing this system is described below.
[0343] The server first analyzes the video data received from the user's device. This analysis uses deep learning libraries such as TensorFlow to extract visual features and simultaneously generate symbolic information. Based on this information, the system automatically tags the data and stores it in storage using MongoDB.
[0344] Next, the server receives a search input from the user. This search query can be received in the form of a prompt, for example, "I'm looking for recommended videos that contain both comedy and heartwarming scenes." The server then uses Elasticsearch to quickly search for relevant video data and generate the best possible search results.
[0345] Furthermore, the server considers viewing history and generates personalized recommendation lists. These lists are displayed by an application built with React Native on the user's device, making it easy for users to find content that matches their interests.
[0346] For example, when a user searches for a comedy movie on the weekend, the system adds relevant movies to a recommendation list based on the user's past viewing and search history, and also presents videos that other viewers have given high ratings to. This allows users to select the content that best suits them from a vast amount of video data.
[0347] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0348] Step 1:
[0349] The user uploads video data to the system using their device. The uploaded video data is sent to the server. This data becomes the input for subsequent processing.
[0350] Step 2:
[0351] The server analyzes the received video data using TensorFlow. The analysis extracts visual features from the video and generates related symbolic information (text) based on these features. This process provides specific tag information for the video data.
[0352] Step 3:
[0353] The server automatically tags the extracted visual features and symbolic information. These tags are stored in MongoDB and output as a tagged database that will serve as the basis for future searches and recommendations.
[0354] Step 4:
[0355] The user enters a search query from their device. The query is received as a prompt that can be processed using natural language processing technology (e.g., "I'm looking for recommended videos that contain both comedy and heartwarming scenes."). This prompt becomes the input to the server.
[0356] Step 5:
[0357] The server uses Elasticsearch to match tag information in the database with search queries and quickly retrieve relevant video data. This process outputs search results for the video that best suits the user's request.
[0358] Step 6:
[0359] The server generates personalized recommendation lists based on viewing history and search queries. To achieve this, it analyzes past viewing activity and selects highly relevant content. This process results in an optimized video recommendation list for the user.
[0360] Step 7:
[0361] The generated search results and recommendation lists are displayed in a device application built with React Native. Users can then select and view video content that interests them based on the information presented.
[0362] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0363] Embodiments of the present invention include a system incorporating a user terminal, a central processing server, a network connecting them, and an emotion engine that recognizes the user's emotional state. This system allows the user to efficiently search for and manage video data corresponding to their emotions.
[0364] When a user selects video data through their device and uploads it to the system, this data is sent to a server. The server analyzes the received data using an AI model and extracts visual features and text information. Based on this information, automatic tagging is performed and stored in a database. This process enables efficient searching of video content.
[0365] Furthermore, the system incorporates an emotion engine that analyzes the user's emotions in real time while they are viewing videos. The emotion engine reads the user's emotional state from their facial expressions and voice and sends this information to the server. The server can then use this emotion information to recommend video content that is appropriate for the user's emotions and to adjust the search results.
[0366] As a concrete example, when a user views "soothing animal videos," the system's emotion engine analyzes the user's facial expressions and reactions. If the user appears relaxed and happy, the server prioritizes displaying similar relaxing video content. This process allows users to easily find video content that suits their emotional state.
[0367] Furthermore, by incorporating analysis information from the emotion engine into the tagging process, the tagging accuracy of video data is improved, resulting in more personalized search results. This makes it possible to provide users with the information they are looking for more quickly and accurately.
[0368] In this way, this system enables flexible management and retrieval of video data that takes into account the user's emotional state, thereby improving the user experience. By having the entire system work together to provide information that meets the user's needs, it can significantly improve efficiency in work and personal activities.
[0369] The following describes the processing flow.
[0370] Step 1:
[0371] The user uses a terminal to select video data and initiate the upload to the system. The terminal prepares the video data and transfers it to the server via the network.
[0372] Step 2:
[0373] The server processes the video data received from the terminal, using an AI model to extract visual features from the data. These extracted features are then used for tagging.
[0374] Step 3:
[0375] The server converts the audio information from the video data into text and analyzes the content within the video. This analysis is performed to generate highly relevant tags.
[0376] Step 4:
[0377] The server integrates visual features and text information and performs automatic tagging. This process assigns keywords related to the video data and stores them in a database.
[0378] Step 5:
[0379] When a user operates the device and begins watching a video, the emotion engine activates and starts analyzing the user's facial expressions and vocal characteristics.
[0380] Step 6:
[0381] The emotion engine analyzes the user's emotional state in real time and sends the obtained emotional information to the server.
[0382] Step 7:
[0383] The server selects video content that matches the user's emotions based on the received emotional information. In this process, videos with tags that match the user's emotions are given priority.
[0384] Step 8:
[0385] The user enters a search query on their device to search for a specific video. The device then sends this query to the server.
[0386] Step 9:
[0387] The server integrates search queries, tagging information, and user sentiment information to retrieve the most suitable video data from the database.
[0388] Step 10:
[0389] The search results are listed, ranked according to the user's emotional state, and sent to the device.
[0390] Step 11:
[0391] The terminal displays ranked search results received from the server to the user, presenting them in a visually easy-to-understand format.
[0392] Step 12:
[0393] If a user needs a summary of a specific video, they can request the server to generate the summary through their device.
[0394] Step 13:
[0395] The server uses generation technology to create a summary of the selected video content and provides it to the terminal.
[0396] Step 14:
[0397] Users can view summaries generated on their devices and quickly grasp the necessary information.
[0398] These steps enable users to efficiently search for and watch videos that resonate with their emotions, and to quickly understand the content they need.
[0399] (Example 2)
[0400] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0401] In today's digital society, it is crucial for users to quickly and accurately find appropriate digital content that matches their individual emotional state. However, conventional systems lacked the means to analyze user emotions in real time and recommend appropriate content, hindering improvements in the user experience. Furthermore, systems capable of dynamic content recommendation based on emotions were limited, and there was a need for information provision tailored to the individual needs of users.
[0402] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0403] In this invention, the server includes means for providing a terminal for users to upload data, means for using machine learning techniques to analyze the received data and extract its characteristics, means for automatically associating and storing identification information with the data, means for analyzing the user's emotional state, means for selecting and providing relevant information based on the emotional state, and means for dynamically adjusting the search results of the data according to the user's emotions. This enables advanced data recommendation and search that takes into account the user's emotional state, resulting in a significant improvement in the user experience.
[0404] "A terminal for users to upload data" refers to an electronic device used to transmit digital data specified by the user to the system.
[0405] "Means of using machine learning technology" refers to technologies that apply artificial intelligence algorithms to analyze received digital data and automatically identify and extract visual and textual information.
[0406] "Means for automatically associating and storing identification information with data" refers to the process of automatically assigning tags and metadata to analyzed digital data and saving that data to a storage device such as a database.
[0407] "Means for analyzing a user's emotional state" refers to technologies that analyze data such as a user's facial expressions and voice to identify the emotions they are expressing.
[0408] "Means of selecting and providing relevant information based on the aforementioned emotional state" refers to a process that automatically selects and presents digital content appropriate to the user's emotions based on the analyzed emotions of that user.
[0409] "Means of dynamically adjusting search results in response to user sentiment" refers to a mechanism that takes user sentiment data into consideration, modifies normal search results in real time, and prioritizes displaying the information most relevant to the user.
[0410] To implement this invention, a system is required in which a user, a terminal, and a server work together. The user can access the system using the terminal and select video data that they wish to view or upload. The terminal has an intuitive interface, making it easy for the user to operate.
[0411] In analyzing video data, the server utilizes machine learning techniques to analyze the received video data. Specifically, it can use widely used deep learning frameworks such as TensorFlow or PyTorch. The server uses these to extract visual features and text information from the video and automatically assigns identification information as tags. This automated tagging process allows the data to be efficiently stored in a database, enabling rapid searching and retrieval later on.
[0412] Furthermore, the server is equipped with an emotion engine to analyze the user's emotional state. This engine analyzes facial expressions and voice data transmitted from the user's device, allowing it to understand the user's emotional state in real time. Facial recognition technology and voice analysis technology are used for the analysis.
[0413] Based on the acquired emotional information, the server recommends video content that is appropriate for the user's current mood. In this recommendation process, the content recommendation algorithm reflects the emotional data and prioritizes providing the most relevant information to the user. For example, if a user is relaxing after watching "soothing animal videos," the server will continue to prioritize displaying videos with similar relaxing effects.
[0414] As a concrete example, here is an example of a prompt statement that utilizes a generative AI model:
[0415] "Please search for relaxing animal videos. If the user smiles, please recommend similar videos."
[0416] This significantly improves the user experience and enables content viewing tailored to individual emotional states. The integrated operation of the entire system makes managing and searching for digital content even more efficient.
[0417] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0418] Step 1:
[0419] The user selects digital data and uploads it to the system from the terminal. The input is a video file selected by the user, which is sent to the server via the terminal's interface. The output is the data arriving at the server. The terminal provides the user with a simple and intuitive UI, making file selection and uploading easy.
[0420] Step 2:
[0421] The server analyzes the received video data. The input is a video file, which is then fed into an AI model to extract visual features and textual information. Deep learning frameworks such as TensorFlow and PyTorch are used for data processing, and automatically tagged data is generated as output. The server efficiently stores this information in a database.
[0422] Step 3:
[0423] The system receives data to analyze the user's emotional state in real time. Input includes facial expressions and voice data sent from the device, which the emotion engine processes to determine the user's emotional state. Emotional information is generated as output and sent to the server. Facial recognition and voice tone analysis technologies are applied in the analysis.
[0424] Step 4:
[0425] The server uses user sentiment information to recommend content. It uses sentiment information and video content from its database as input. A content recommendation algorithm selects videos that match the sentiment, generating a list of candidates to display to the user. The server then sends this list to the user's device.
[0426] Step 5:
[0427] The user views recommended video content on their device. The input is a list of content sent from the server, from which the user can select and begin playback. The output further collects the user's emotional state and sends it back to the system as feedback to help improve future recommendations. This allows for continuous improvement of the user experience.
[0428] (Application Example 2)
[0429] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".
[0430] Modern video streaming services are required to appropriately select and efficiently deliver the video content that users want, but content recommendations that take into account the user's emotional state are not being adequately implemented. As a result, user satisfaction is low, and the value of video streaming services is not being fully realized.
[0431] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0432] In this invention, the server includes means for analyzing the user's facial expression and voice information to determine their emotional state, means for recommending video content based on the determined emotional state, and means for providing search results to the user's terminal. This makes it possible to recommend video content based on the user's emotional state, enabling the efficient and effective provision of the video content the user desires.
[0433] "Video data" refers to digital data containing visual information, such as videos and images.
[0434] "Means of receiving" refers to mechanisms and methods for taking in data from an external source, such as acquiring information via networks or communication devices.
[0435] "Methods for analysis and extraction of visual and textual features" refers to methods for analyzing the content of video data and extracting features from image and textual information.
[0436] "Methods for automatic tagging" refer to the process of automatically assigning identifiable words or labels to data based on extracted features.
[0437] "Means of storing information in a database" refers to a system for long-term storage of structured information.
[0438] "A means of receiving search queries and searching for relevant video data" refers to a method of taking in information requests from users and finding corresponding video data.
[0439] "Means of providing to the user's terminal" refers to functions for displaying or delivering search results and recommended content to the user's device.
[0440] "A means of analyzing a user's facial expression and voice information to determine their emotional state" refers to a process of identifying a user's emotions at a given time by analyzing their facial movements and tone of voice.
[0441] "Methods for recommending video content based on emotional state" refer to technologies that take into account the determined emotions and select and present appropriate videos.
[0442] This system is implemented through a network that includes the user's smart device and a server. Video data viewed by the user is transmitted to the server in real time via a device such as a smartphone. The server analyzes the received video data and extracts visual features and text information. Feature extraction is performed using deep learning libraries such as OpenCV and TensorFlow.
[0443] Based on the analyzed features, a generative AI model automatically tags the data and stores it in a structured database. Furthermore, software equipped with an emotion engine detects changes in the user's facial expressions and voice to determine their emotional state. This analysis utilizes the smartphone's hardware, including its camera and microphone.
[0444] Once the user's emotional state is determined, the system uses that information to recommend video content that best suits the user's current emotions. This makes it easy for users to find content that resonates with their feelings.
[0445] For example, if the system detects that a user is crying while watching an emotionally moving film, it will recommend another film that is both touching and heartwarming. If the system detects laughter in the user's voice, it will then display a more entertaining comedy.
[0446] Here are some examples of prompts generated using a generative AI model. Instructions such as "If the user smiles, please display the following recommended content" or "If the user shows an emotion such as crying, please recommend a heartwarming movie" are used.
[0447] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0448] Step 1:
[0449] The user selects the video content they want to watch via their smart device. The device receives the video data and sends it to the server in real time. In this step, the video file is the input, and the output is data sent to the server in streaming format.
[0450] Step 2:
[0451] The server uses an AI model to extract visual features and text information to analyze the received video data. The input is video data, and the output is a list of features. Image analysis and natural language processing are performed using libraries such as OpenCV and TensorFlow, and the features are extracted as basic information for automatic tagging.
[0452] Step 3:
[0453] The server automatically tags the extracted visual and text features using a generative AI model. This tagging information is stored in a database. The input is a list of features, and the output is a data structure with automatic tags. The database storage process takes place at this stage.
[0454] Step 4:
[0455] When a user views video using their device, the device uses its camera and microphone to collect facial and audio information in real time. The collected data is sent to a server. The input consists of facial and audio data, and the output is the transmission of a data stream to the server.
[0456] Step 5:
[0457] The server analyzes the received facial expression and voice information to determine the user's emotional state. The input consists of facial expression data and voice data, and the output is a determination result indicating the emotional state (e.g., joy, sadness, etc.). The emotion engine performs this data calculation.
[0458] Step 6:
[0459] The server recommends video content that matches the user's current mood based on the determined emotional state. The input is the emotional state determination result and video database information, and the output is a list of recommended video content. A generative AI model is used for this recommendation process.
[0460] Step 7:
[0461] The terminal provides the user with video content recommended by the server. The user selects the content they want to watch next from the recommendation list displayed on the terminal. The input is the recommendation list received from the server, and the output is the video content selected by the user.
[0462] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0463] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0464] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0465] [Third Embodiment]
[0466] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0467] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0468] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0469] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0470] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0471] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0472] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0473] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0474] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0475] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0476] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0477] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0478] Embodiments of the present invention are systems comprising a user-operated terminal, a central processing server, and a network that connects these. This system automatically analyzes and tags video data provided by the user, streamlining subsequent searching and management.
[0479] When a user uploads video data to the system via their device, the server receives the data. The server uses an AI model to analyze the visual features of the video and extract related text information. This analysis automatically tags the video, and the information is stored in a database.
[0480] This tagged data allows users to enter search queries from their devices and quickly find corresponding video content. The search is performed by the server associating the tag information of relevant videos with the search query, providing the user with the most suitable results.
[0481] For example, if a user uploads video data to the system that includes "a family's daily life filmed at the zoo," the server automatically generates tags such as "animals," "family," "children," and "outdoors" from this video and registers them in the database. Later, when the user performs a search using the query "children and animals," the server quickly searches for relevant videos and provides the results to the terminal.
[0482] Furthermore, this system uses generation technology to create summaries of selected videos, helping users effectively manage content. These summaries include important scenes and key phrases from the video, allowing users to instantly grasp the content.
[0483] These features allow users to efficiently manage vast amounts of video data and quickly access the information they need. Furthermore, this system reduces the burden of data management, contributing to improved efficiency in work and daily life.
[0484] The following describes the processing flow.
[0485] Step 1:
[0486] The user uses their device to select video data and perform the upload operation. At that time, preparations are made to send the data to the server via the internet.
[0487] Step 2:
[0488] The device sends the video data to be uploaded to the server. Once the data transmission is complete, the user is notified that the transmission is complete.
[0489] Step 3:
[0490] The server then processes the received video data. First, the data is input into an AI model, and visual features are extracted using image recognition technology.
[0491] Step 4:
[0492] The server uses speech recognition technology to extract text information from the audio of the video. This text information is useful for data retrieval in later processes.
[0493] Step 5:
[0494] The server integrates extracted visual features and text information using multiple modal learning technologies to gain a deeper understanding of the meaning of the video content.
[0495] Step 6:
[0496] The server automatically generates relevant tags based on the information obtained through analysis. This tagging will improve future search performance.
[0497] Step 7:
[0498] The server saves the tagged video data to a database, making it searchable. Once saving is complete, the results are recorded as a log.
[0499] Step 8:
[0500] The user enters a search query through their device and requests to search for a specific video. The device then sends this search request to the server.
[0501] Step 9:
[0502] The server analyzes the received search query and searches the database for video data related to the query. During the matching process, it understands the meaning of the search query and selects the most suitable search results based on tag information.
[0503] Step 10:
[0504] The server lists the search results and sends them back to the terminal. The user then reviews the results visually in an easy-to-understand format on the terminal.
[0505] Step 11:
[0506] If a user requires more detailed information, they can request a summary of a specific video through their device.
[0507] Step 12:
[0508] The server uses generation technology to automatically generate a summary of the specified video content and sends that content to the terminal.
[0509] Step 13:
[0510] Users can view the summary generated on their device and quickly understand the content.
[0511] These steps enable users to efficiently search and manage video data and quickly access the information they need.
[0512] (Example 1)
[0513] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0514] There is a need to efficiently analyze vast amounts of video data and quickly search for and provide relevant information. However, conventional methods require significant time and effort for manual tagging and searching, making efficient data management and information access difficult. Furthermore, extracting necessary information requires advanced analysis and summarization, which needs to be automated.
[0515] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0516] In this invention, the server includes a device for receiving video data, a device for analyzing the received video data and extracting visual attributes and text information, and a device for automatically tagging based on those attributes. This automates the analysis of vast amounts of video data, enabling efficient tagging and searching.
[0517] "Video data" refers to a series of images or video data containing visual information, which are recorded or transmitted electronically.
[0518] "Device" refers to a system or part of a hardware or software designed to perform a specific function or process.
[0519] "Analysis" refers to the process of examining data or information in detail to understand its structure and meaning.
[0520] "Visual attributes" refer to visually-based features such as objects, shapes, and colors present within video data.
[0521] "Text information" refers to information consisting of characters and words related to video data, including descriptions and tags extracted from visual elements.
[0522] "Automatic tagging" refers to the process of mechanically adding keywords and categories related to video data using AI or algorithms.
[0523] "Information recording media" refers to physical or virtual storage used to store and maintain digital data over long periods of time.
[0524] A "search request" refers to a query or question entered by a user to retrieve specific information.
[0525] A "user-operated device" refers to a computer or portable device that a human user directly operates and uses to interact with a system.
[0526] "Generative technologies" refer to a set of technical methods or tools for generating automatically generated data and information.
[0527] Embodiments of this invention include a system comprising a user-operated terminal, a central data processing server, and a network connecting them. The aim of this system is to efficiently manage vast amounts of video data and enable rapid searching and access to information.
[0528] The user operates a terminal to upload video data to the system. The terminal sends the data to the server via the internet. The server inputs the video data into an AI model using Python (e.g., TensorFlow or PyTorch) and performs extraction of visual attributes and text information. This AI model performs image analysis on each video frame, generating tags by detecting objects and recognizing scenes.
[0529] The server automatically tags the video data using the generated tags and stores it in an information storage medium, such as a NoSQL database (e.g., MongoDB). This stored data can be quickly accessed in response to subsequent search requests.
[0530] The user sends a search request using keywords from their device to the server. The server executes the search request against the database and generates a list of video data that matches the query. This result is sent back to the user's device, and the user can view the results directly.
[0531] Furthermore, the server uses generation technology to generate a summary of the selected video data. This summary utilizes a natural language processing model (e.g., using Hugging Face's Transformers) to transcribe important scenes and key phrases from the video into text, allowing the user to quickly grasp the content.
[0532] For example, if a user uploads "a video of their family taken at a zoo," the server generates tags such as "animals," "family," "children," and "outdoors." This tag information is stored in a database and used later when a search request for "children and animals" is received. The server quickly searches for relevant videos and provides the results to the user's device.
[0533] An example of a prompt message is, "Analyze the following video and generate relevant tags. Based on the features in the video, prioritize extracting tags related to family or animals." This prompt is used as input to the AI model, which helps in detailed tag generation and efficient data management.
[0534] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0535] Step 1:
[0536] The user uses their device to select and upload the video data they want to analyze. The device then takes the video data as input and sends it to the server via the network. During this process, a progress bar is displayed to show the status of data transmission, giving the user peace of mind.
[0537] Step 2:
[0538] The server receives video data transmitted from the terminal and performs a hash check to ensure data integrity. The received video data becomes the input for the analysis process. The server feeds the video data into an AI model (e.g., using TensorFlow or PyTorch) and extracts visual features for each image frame. This process performs object detection and scene recognition, and outputs extracted visual attributes and text information.
[0539] Step 3:
[0540] The server automatically tags the video based on visual features extracted by the AI model. The generated tags indicate the content of the video, such as "animals" or "children." These tags are added to the video data, resulting in tagged data output. The server stores this tagged data on an information storage medium. A NoSQL database (e.g., MongoDB) is used for storage to ensure fast access and scalability.
[0541] Step 4:
[0542] The user enters a keyword-based search request from their device and sends it to the server. The search query arrives at the server as input. The server parses the query and matches it with tag information in the database. This process extracts matching video data and outputs a list of search results. This list is then sent back to the user's device.
[0543] Step 5:
[0544] The server uses generative techniques to create a summary from the video data selected by the user. A natural language processing model (e.g., using Hugging Face's Transformers) extracts important scenes and key phrases from the video and provides output as a text summary. The server sends this summary information to the user's terminal, allowing the user to easily understand the content of the video.
[0545] (Application Example 1)
[0546] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0547] In today's world, with the ever-increasing volume of video content, users are required to quickly search for and view the information they desire from a vast amount of video data. However, manual tagging and classification are labor-intensive, making efficient management difficult. Furthermore, the lack of personalized recommendation features means users may miss out on content that interests them. It is necessary to solve these problems and provide an environment where users can enjoy video content efficiently and accurately.
[0548] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0549] In this invention, the server includes a device for extracting visual features and symbolic information, a device for automatically tagging content, and a device for providing personalized recommendations using viewing history. This enables users to efficiently manage vast amounts of video content and receive personalized recommendations based on visual features and viewing history.
[0550] "Video data" refers to digital information that records visual information continuously over time.
[0551] A "device" is a set of hardware or software designed to perform a specific function.
[0552] "Analysis" is the process of breaking down data and information into detailed components and understanding and evaluating their elements and structure.
[0553] "Visual features" are attributes that are identified by their appearance, such as color, shape, and pattern, contained within an image or video.
[0554] "Symbolic information" refers to text data that extracts specific meanings and content from video data and explains them.
[0555] "Tagging" is the process of assigning identifying labels or keywords to data.
[0556] A "storage device" is a medium or device capable of storing information and data for a long period of time.
[0557] "Search input" refers to the provision of queries or keywords that a user makes to the system in order to find specific data.
[0558] An "information processing system" is a set of computer systems configured to collect, store, analyze, and process data in order to provide necessary information.
[0559] A "recommendation list" is a list that presents highly relevant options based on the user's interests and past behavior.
[0560] "Personalized recommendations" are a process that presents the most suitable options based on the user's specific interests and preferences.
[0561] The system for implementing this invention is designed to efficiently manage video data and provide that data to users in a useful format. The method for implementing this system is described below.
[0562] The server first analyzes the video data received from the user's device. This analysis uses deep learning libraries such as TensorFlow to extract visual features and simultaneously generate symbolic information. Based on this information, the system automatically tags the data and stores it in storage using MongoDB.
[0563] Next, the server receives a search input from the user. This search query can be received in the form of a prompt, for example, "I'm looking for recommended videos that contain both comedy and heartwarming scenes." The server then uses Elasticsearch to quickly search for relevant video data and generate the best possible search results.
[0564] Furthermore, the server considers viewing history and generates personalized recommendation lists. These lists are displayed by an application built with React Native on the user's device, making it easy for users to find content that matches their interests.
[0565] For example, when a user searches for a comedy movie on the weekend, the system adds relevant movies to a recommendation list based on the user's past viewing and search history, and also presents videos that other viewers have given high ratings to. This allows users to select the content that best suits them from a vast amount of video data.
[0566] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0567] Step 1:
[0568] The user uploads video data to the system using their device. The uploaded video data is sent to the server. This data becomes the input for subsequent processing.
[0569] Step 2:
[0570] The server analyzes the received video data using TensorFlow. The analysis extracts visual features from the video and generates related symbolic information (text) based on these features. This process provides specific tag information for the video data.
[0571] Step 3:
[0572] The server automatically tags the extracted visual features and symbolic information. These tags are stored in MongoDB and output as a tagged database that will serve as the basis for future searches and recommendations.
[0573] Step 4:
[0574] The user enters a search query from their device. The query is received as a prompt that can be processed using natural language processing technology (e.g., "I'm looking for recommended videos that contain both comedy and heartwarming scenes."). This prompt becomes the input to the server.
[0575] Step 5:
[0576] The server uses Elasticsearch to match tag information in the database with search queries and quickly retrieve relevant video data. This process outputs search results for the video that best suits the user's request.
[0577] Step 6:
[0578] The server generates personalized recommendation lists based on viewing history and search queries. To achieve this, it analyzes past viewing activity and selects highly relevant content. This process results in an optimized video recommendation list for the user.
[0579] Step 7:
[0580] The generated search results and recommendation lists are displayed in a device application built with React Native. Users can then select and view video content that interests them based on the information presented.
[0581] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0582] Embodiments of the present invention include a system incorporating a user terminal, a central processing server, a network connecting them, and an emotion engine that recognizes the user's emotional state. This system allows the user to efficiently search for and manage video data corresponding to their emotions.
[0583] When a user selects video data through their device and uploads it to the system, this data is sent to a server. The server analyzes the received data using an AI model and extracts visual features and text information. Based on this information, automatic tagging is performed and stored in a database. This process enables efficient searching of video content.
[0584] Furthermore, the system incorporates an emotion engine that analyzes the user's emotions in real time while they are viewing videos. The emotion engine reads the user's emotional state from their facial expressions and voice and sends this information to the server. The server can then use this emotion information to recommend video content that is appropriate for the user's emotions and to adjust the search results.
[0585] As a concrete example, when a user views "soothing animal videos," the system's emotion engine analyzes the user's facial expressions and reactions. If the user appears relaxed and happy, the server prioritizes displaying similar relaxing video content. This process allows users to easily find video content that suits their emotional state.
[0586] Furthermore, by incorporating analysis information from the emotion engine into the tagging process, the tagging accuracy of video data is improved, resulting in more personalized search results. This makes it possible to provide users with the information they are looking for more quickly and accurately.
[0587] In this way, this system enables flexible management and retrieval of video data that takes into account the user's emotional state, thereby improving the user experience. By having the entire system work together to provide information that meets the user's needs, it can significantly improve efficiency in work and personal activities.
[0588] The following describes the processing flow.
[0589] Step 1:
[0590] The user uses a terminal to select video data and initiate the upload to the system. The terminal prepares the video data and transfers it to the server via the network.
[0591] Step 2:
[0592] The server processes the video data received from the terminal, using an AI model to extract visual features from the data. These extracted features are then used for tagging.
[0593] Step 3:
[0594] The server converts the audio information from the video data into text and analyzes the content within the video. This analysis is performed to generate highly relevant tags.
[0595] Step 4:
[0596] The server integrates visual features and text information and performs automatic tagging. This process assigns keywords related to the video data and stores them in a database.
[0597] Step 5:
[0598] When a user operates the device and begins watching a video, the emotion engine activates and starts analyzing the user's facial expressions and vocal characteristics.
[0599] Step 6:
[0600] The emotion engine analyzes the user's emotional state in real time and sends the obtained emotional information to the server.
[0601] Step 7:
[0602] The server selects video content that matches the user's emotions based on the received emotional information. In this process, videos with tags that match the user's emotions are given priority.
[0603] Step 8:
[0604] The user enters a search query on their device to search for a specific video. The device then sends this query to the server.
[0605] Step 9:
[0606] The server integrates search queries, tagging information, and user sentiment information to retrieve the most suitable video data from the database.
[0607] Step 10:
[0608] The search results are listed, ranked according to the user's emotional state, and sent to the device.
[0609] Step 11:
[0610] The terminal displays ranked search results received from the server to the user, presenting them in a visually easy-to-understand format.
[0611] Step 12:
[0612] If a user needs a summary of a specific video, they can request the server to generate the summary through their device.
[0613] Step 13:
[0614] The server uses generation technology to create a summary of the selected video content and provides it to the terminal.
[0615] Step 14:
[0616] Users can view summaries generated on their devices and quickly grasp the necessary information.
[0617] These steps enable users to efficiently search for and watch videos that resonate with their emotions, and to quickly understand the content they need.
[0618] (Example 2)
[0619] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0620] In today's digital society, it is crucial for users to quickly and accurately find appropriate digital content that matches their individual emotional state. However, conventional systems lacked the means to analyze user emotions in real time and recommend appropriate content, hindering improvements in the user experience. Furthermore, systems capable of dynamic content recommendation based on emotions were limited, and there was a need for information provision tailored to the individual needs of users.
[0621] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0622] In this invention, the server includes means for providing a terminal for users to upload data, means for using machine learning techniques to analyze the received data and extract its characteristics, means for automatically associating and storing identification information with the data, means for analyzing the user's emotional state, means for selecting and providing relevant information based on the emotional state, and means for dynamically adjusting the search results of the data according to the user's emotions. This enables advanced data recommendation and search that takes into account the user's emotional state, resulting in a significant improvement in the user experience.
[0623] "A terminal for users to upload data" refers to an electronic device used to transmit digital data specified by the user to the system.
[0624] "Means of using machine learning technology" refers to technologies that apply artificial intelligence algorithms to analyze received digital data and automatically identify and extract visual and textual information.
[0625] "Means for automatically associating and storing identification information with data" refers to the process of automatically assigning tags and metadata to analyzed digital data and saving that data to a storage device such as a database.
[0626] "Means for analyzing a user's emotional state" refers to technologies that analyze data such as a user's facial expressions and voice to identify the emotions they are expressing.
[0627] "Means of selecting and providing relevant information based on the aforementioned emotional state" refers to a process that automatically selects and presents digital content appropriate to the user's emotions based on the analyzed emotions of that user.
[0628] "Means of dynamically adjusting search results in response to user sentiment" refers to a mechanism that takes user sentiment data into consideration, modifies normal search results in real time, and prioritizes displaying the information most relevant to the user.
[0629] To implement this invention, a system is required in which a user, a terminal, and a server work together. The user can access the system using the terminal and select video data that they wish to view or upload. The terminal has an intuitive interface, making it easy for the user to operate.
[0630] In analyzing video data, the server utilizes machine learning techniques to analyze the received video data. Specifically, it can use widely used deep learning frameworks such as TensorFlow or PyTorch. The server uses these to extract visual features and text information from the video and automatically assigns identification information as tags. This automated tagging process allows the data to be efficiently stored in a database, enabling rapid searching and retrieval later on.
[0631] Furthermore, the server is equipped with an emotion engine to analyze the user's emotional state. This engine analyzes facial expressions and voice data transmitted from the user's device, allowing it to understand the user's emotional state in real time. Facial recognition technology and voice analysis technology are used for the analysis.
[0632] Based on the acquired emotional information, the server recommends video content that is appropriate for the user's current mood. In this recommendation process, the content recommendation algorithm reflects the emotional data and prioritizes providing the most relevant information to the user. For example, if a user is relaxing after watching "soothing animal videos," the server will continue to prioritize displaying videos with similar relaxing effects.
[0633] As a concrete example, here is an example of a prompt statement that utilizes a generative AI model:
[0634] "Please search for relaxing animal videos. If the user smiles, please recommend similar videos."
[0635] This significantly improves the user experience and enables content viewing tailored to individual emotional states. The integrated operation of the entire system makes managing and searching for digital content even more efficient.
[0636] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0637] Step 1:
[0638] The user selects digital data and uploads it to the system from the terminal. The input is a video file selected by the user, which is sent to the server via the terminal's interface. The output is the data arriving at the server. The terminal provides the user with a simple and intuitive UI, making file selection and uploading easy.
[0639] Step 2:
[0640] The server analyzes the received video data. The input is a video file, which is then fed into an AI model to extract visual features and textual information. Deep learning frameworks such as TensorFlow and PyTorch are used for data processing, and automatically tagged data is generated as output. The server efficiently stores this information in a database.
[0641] Step 3:
[0642] The system receives data to analyze the user's emotional state in real time. Input includes facial expressions and voice data sent from the device, which the emotion engine processes to determine the user's emotional state. Emotional information is generated as output and sent to the server. Facial recognition and voice tone analysis technologies are applied in the analysis.
[0643] Step 4:
[0644] The server uses user sentiment information to recommend content. It uses sentiment information and video content from its database as input. A content recommendation algorithm selects videos that match the sentiment, generating a list of candidates to display to the user. The server then sends this list to the user's device.
[0645] Step 5:
[0646] The user views recommended video content on their device. The input is a list of content sent from the server, from which the user can select and begin playback. The output further collects the user's emotional state and sends it back to the system as feedback to help improve future recommendations. This allows for continuous improvement of the user experience.
[0647] (Application Example 2)
[0648] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0649] Modern video streaming services are required to appropriately select and efficiently deliver the video content that users want, but content recommendations that take into account the user's emotional state are not being adequately implemented. As a result, user satisfaction is low, and the value of video streaming services is not being fully realized.
[0650] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0651] In this invention, the server includes means for analyzing the user's facial expression and voice information to determine their emotional state, means for recommending video content based on the determined emotional state, and means for providing search results to the user's terminal. This makes it possible to recommend video content based on the user's emotional state, enabling the efficient and effective provision of the video content the user desires.
[0652] "Video data" refers to digital data containing visual information, such as videos and images.
[0653] "Means of receiving" refers to mechanisms and methods for taking in data from an external source, such as acquiring information via networks or communication devices.
[0654] "Methods for analysis and extraction of visual and textual features" refers to methods for analyzing the content of video data and extracting features from image and textual information.
[0655] "Methods for automatic tagging" refer to the process of automatically assigning identifiable words or labels to data based on extracted features.
[0656] "Means of storing information in a database" refers to a system for long-term storage of structured information.
[0657] "A means of receiving search queries and searching for relevant video data" refers to a method of taking in information requests from users and finding corresponding video data.
[0658] "Means of providing to the user's terminal" refers to functions for displaying or delivering search results and recommended content to the user's device.
[0659] "A means of analyzing a user's facial expression and voice information to determine their emotional state" refers to a process of identifying a user's emotions at a given time by analyzing their facial movements and tone of voice.
[0660] "Methods for recommending video content based on emotional state" refer to technologies that take into account the determined emotions and select and present appropriate videos.
[0661] This system is implemented through a network that includes the user's smart device and a server. Video data viewed by the user is transmitted to the server in real time via a device such as a smartphone. The server analyzes the received video data and extracts visual features and text information. Feature extraction is performed using deep learning libraries such as OpenCV and TensorFlow.
[0662] Based on the analyzed features, a generative AI model automatically tags the data and stores it in a structured database. Furthermore, software equipped with an emotion engine detects changes in the user's facial expressions and voice to determine their emotional state. This analysis utilizes the smartphone's hardware, including its camera and microphone.
[0663] Once the user's emotional state is determined, the system uses that information to recommend video content that best suits the user's current emotions. This makes it easy for users to find content that resonates with their feelings.
[0664] For example, if the system detects that a user is crying while watching an emotionally moving film, it will recommend another film that is both touching and heartwarming. If the system detects laughter in the user's voice, it will then display a more entertaining comedy.
[0665] Here are some examples of prompts generated using a generative AI model. Instructions such as "If the user smiles, please display the following recommended content" or "If the user shows an emotion such as crying, please recommend a heartwarming movie" are used.
[0666] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0667] Step 1:
[0668] The user selects the video content they want to watch via their smart device. The device receives the video data and sends it to the server in real time. In this step, the video file is the input, and the output is data sent to the server in streaming format.
[0669] Step 2:
[0670] The server uses an AI model to extract visual features and text information to analyze the received video data. The input is video data, and the output is a list of features. Image analysis and natural language processing are performed using libraries such as OpenCV and TensorFlow, and the features are extracted as basic information for automatic tagging.
[0671] Step 3:
[0672] The server automatically tags the extracted visual and text features using a generative AI model. This tagging information is stored in a database. The input is a list of features, and the output is a data structure with automatic tags. The database storage process takes place at this stage.
[0673] Step 4:
[0674] When a user views video using their device, the device uses its camera and microphone to collect facial and audio information in real time. The collected data is sent to a server. The input consists of facial and audio data, and the output is the transmission of a data stream to the server.
[0675] Step 5:
[0676] The server analyzes the received facial expression and voice information to determine the user's emotional state. The input consists of facial expression data and voice data, and the output is a determination result indicating the emotional state (e.g., joy, sadness, etc.). The emotion engine performs this data calculation.
[0677] Step 6:
[0678] The server recommends video content that matches the user's current mood based on the determined emotional state. The input is the emotional state determination result and video database information, and the output is a list of recommended video content. A generative AI model is used for this recommendation process.
[0679] Step 7:
[0680] The terminal provides the user with video content recommended by the server. The user selects the content they want to watch next from the recommendation list displayed on the terminal. The input is the recommendation list received from the server, and the output is the video content selected by the user.
[0681] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0682] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0683] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0684] [Fourth Embodiment]
[0685] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0686] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0687] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0688] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0689] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0690] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0691] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0692] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0693] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0694] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0695] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0696] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0697] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0698] Embodiments of the present invention are systems comprising a user-operated terminal, a central processing server, and a network that connects these. This system automatically analyzes and tags video data provided by the user, streamlining subsequent searching and management.
[0699] When a user uploads video data to the system via their device, the server receives the data. The server uses an AI model to analyze the visual features of the video and extract related text information. This analysis automatically tags the video, and the information is stored in a database.
[0700] This tagged data allows users to enter search queries from their devices and quickly find corresponding video content. The search is performed by the server associating the tag information of relevant videos with the search query, providing the user with the most suitable results.
[0701] For example, if a user uploads video data to the system that includes "a family's daily life filmed at the zoo," the server automatically generates tags such as "animals," "family," "children," and "outdoors" from this video and registers them in the database. Later, when the user performs a search using the query "children and animals," the server quickly searches for relevant videos and provides the results to the terminal.
[0702] Furthermore, this system uses generation technology to create summaries of selected videos, helping users effectively manage content. These summaries include important scenes and key phrases from the video, allowing users to instantly grasp the content.
[0703] These features allow users to efficiently manage vast amounts of video data and quickly access the information they need. Furthermore, this system reduces the burden of data management, contributing to improved efficiency in work and daily life.
[0704] The following describes the processing flow.
[0705] Step 1:
[0706] The user uses their device to select video data and perform the upload operation. At that time, preparations are made to send the data to the server via the internet.
[0707] Step 2:
[0708] The device sends the video data to be uploaded to the server. Once the data transmission is complete, the user is notified that the transmission is complete.
[0709] Step 3:
[0710] The server then processes the received video data. First, the data is input into an AI model, and visual features are extracted using image recognition technology.
[0711] Step 4:
[0712] The server uses speech recognition technology to extract text information from the audio of the video. This text information is useful for data retrieval in later processes.
[0713] Step 5:
[0714] The server integrates extracted visual features and text information using multiple modal learning technologies to gain a deeper understanding of the meaning of the video content.
[0715] Step 6:
[0716] The server automatically generates relevant tags based on the information obtained through analysis. This tagging will improve future search performance.
[0717] Step 7:
[0718] The server saves the tagged video data to a database, making it searchable. Once saving is complete, the results are recorded as a log.
[0719] Step 8:
[0720] The user enters a search query through their device and requests to search for a specific video. The device then sends this search request to the server.
[0721] Step 9:
[0722] The server analyzes the received search query and searches the database for video data related to the query. During the matching process, it understands the meaning of the search query and selects the most suitable search results based on tag information.
[0723] Step 10:
[0724] The server lists the search results and sends them back to the terminal. The user then reviews the results visually in an easy-to-understand format on the terminal.
[0725] Step 11:
[0726] If a user requires more detailed information, they can request a summary of a specific video through their device.
[0727] Step 12:
[0728] The server uses generation technology to automatically generate a summary of the specified video content and sends that content to the terminal.
[0729] Step 13:
[0730] Users can view the summary generated on their device and quickly understand the content.
[0731] These steps enable users to efficiently search and manage video data and quickly access the information they need.
[0732] (Example 1)
[0733] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0734] There is a need to efficiently analyze vast amounts of video data and quickly search for and provide relevant information. However, conventional methods require significant time and effort for manual tagging and searching, making efficient data management and information access difficult. Furthermore, extracting necessary information requires advanced analysis and summarization, which needs to be automated.
[0735] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0736] In this invention, the server includes a device for receiving video data, a device for analyzing the received video data and extracting visual attributes and text information, and a device for automatically tagging based on those attributes. This automates the analysis of vast amounts of video data, enabling efficient tagging and searching.
[0737] "Video data" refers to a series of images or video data containing visual information, which are recorded or transmitted electronically.
[0738] "Device" refers to a system or part of a hardware or software designed to perform a specific function or process.
[0739] "Analysis" refers to the process of examining data or information in detail to understand its structure and meaning.
[0740] "Visual attributes" refer to visually-based features such as objects, shapes, and colors present within video data.
[0741] "Text information" refers to information consisting of characters and words related to video data, including descriptions and tags extracted from visual elements.
[0742] "Automatic tagging" refers to the process of mechanically adding keywords and categories related to video data using AI or algorithms.
[0743] "Information recording media" refers to physical or virtual storage used to store and maintain digital data over long periods of time.
[0744] A "search request" refers to a query or question entered by a user to retrieve specific information.
[0745] A "user-operated device" refers to a computer or portable device that a human user directly operates and uses to interact with a system.
[0746] "Generative technologies" refer to a set of technical methods or tools for generating automatically generated data and information.
[0747] Embodiments of this invention include a system comprising a user-operated terminal, a central data processing server, and a network connecting them. The aim of this system is to efficiently manage vast amounts of video data and enable rapid searching and access to information.
[0748] The user operates a terminal to upload video data to the system. The terminal sends the data to the server via the internet. The server inputs the video data into an AI model using Python (e.g., TensorFlow or PyTorch) and performs extraction of visual attributes and text information. This AI model performs image analysis on each video frame, generating tags by detecting objects and recognizing scenes.
[0749] The server automatically tags the video data using the generated tags and stores it in an information storage medium, such as a NoSQL database (e.g., MongoDB). This stored data can be quickly accessed in response to subsequent search requests.
[0750] The user sends a search request using keywords from their device to the server. The server executes the search request against the database and generates a list of video data that matches the query. This result is sent back to the user's device, and the user can view the results directly.
[0751] Furthermore, the server uses generation technology to generate a summary of the selected video data. This summary utilizes a natural language processing model (e.g., using Hugging Face's Transformers) to transcribe important scenes and key phrases from the video into text, allowing the user to quickly grasp the content.
[0752] For example, if a user uploads "a video of their family taken at a zoo," the server generates tags such as "animals," "family," "children," and "outdoors." This tag information is stored in a database and used later when a search request for "children and animals" is received. The server quickly searches for relevant videos and provides the results to the user's device.
[0753] An example of a prompt message is, "Analyze the following video and generate relevant tags. Based on the features in the video, prioritize extracting tags related to family or animals." This prompt is used as input to the AI model, which helps in detailed tag generation and efficient data management.
[0754] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0755] Step 1:
[0756] The user uses their device to select and upload the video data they want to analyze. The device then takes the video data as input and sends it to the server via the network. During this process, a progress bar is displayed to show the status of data transmission, giving the user peace of mind.
[0757] Step 2:
[0758] The server receives video data transmitted from the terminal and performs a hash check to ensure data integrity. The received video data becomes the input for the analysis process. The server feeds the video data into an AI model (e.g., using TensorFlow or PyTorch) and extracts visual features for each image frame. This process performs object detection and scene recognition, and outputs extracted visual attributes and text information.
[0759] Step 3:
[0760] The server automatically tags the video based on visual features extracted by the AI model. The generated tags indicate the content of the video, such as "animals" or "children." These tags are added to the video data, resulting in tagged data output. The server stores this tagged data on an information storage medium. A NoSQL database (e.g., MongoDB) is used for storage to ensure fast access and scalability.
[0761] Step 4:
[0762] The user enters a keyword-based search request from their device and sends it to the server. The search query arrives at the server as input. The server parses the query and matches it with tag information in the database. This process extracts matching video data and outputs a list of search results. This list is then sent back to the user's device.
[0763] Step 5:
[0764] The server uses generative techniques to create a summary from the video data selected by the user. A natural language processing model (e.g., using Hugging Face's Transformers) extracts important scenes and key phrases from the video and provides output as a text summary. The server sends this summary information to the user's terminal, allowing the user to easily understand the content of the video.
[0765] (Application Example 1)
[0766] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0767] In today's world, with the ever-increasing volume of video content, users are required to quickly search for and view the information they desire from a vast amount of video data. However, manual tagging and classification are labor-intensive, making efficient management difficult. Furthermore, the lack of personalized recommendation features means users may miss out on content that interests them. It is necessary to solve these problems and provide an environment where users can enjoy video content efficiently and accurately.
[0768] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0769] In this invention, the server includes a device for extracting visual features and symbolic information, a device for automatically tagging content, and a device for providing personalized recommendations using viewing history. This enables users to efficiently manage vast amounts of video content and receive personalized recommendations based on visual features and viewing history.
[0770] "Video data" refers to digital information that records visual information continuously over time.
[0771] A "device" is a set of hardware or software designed to perform a specific function.
[0772] "Analysis" is the process of breaking down data and information into detailed components and understanding and evaluating their elements and structure.
[0773] "Visual features" are attributes that are identified by their appearance, such as color, shape, and pattern, contained within an image or video.
[0774] "Symbolic information" refers to text data that extracts specific meanings and content from video data and explains them.
[0775] "Tagging" is the process of assigning identifying labels or keywords to data.
[0776] A "storage device" is a medium or device capable of storing information and data for a long period of time.
[0777] "Search input" refers to the provision of queries or keywords that a user makes to the system in order to find specific data.
[0778] An "information processing system" is a set of computer systems configured to collect, store, analyze, and process data in order to provide necessary information.
[0779] A "recommendation list" is a list that presents highly relevant options based on the user's interests and past behavior.
[0780] "Personalized recommendations" are a process that presents the most suitable options based on the user's specific interests and preferences.
[0781] The system for implementing this invention is designed to efficiently manage video data and provide that data to users in a useful format. The method for implementing this system is described below.
[0782] The server first analyzes the video data received from the user's device. This analysis uses deep learning libraries such as TensorFlow to extract visual features and simultaneously generate symbolic information. Based on this information, the system automatically tags the data and stores it in storage using MongoDB.
[0783] Next, the server receives a search input from the user. This search query can be received in the form of a prompt, for example, "I'm looking for recommended videos that contain both comedy and heartwarming scenes." The server then uses Elasticsearch to quickly search for relevant video data and generate the best possible search results.
[0784] Furthermore, the server considers viewing history and generates personalized recommendation lists. These lists are displayed by an application built with React Native on the user's device, making it easy for users to find content that matches their interests.
[0785] For example, when a user searches for a comedy movie on the weekend, the system adds relevant movies to a recommendation list based on the user's past viewing and search history, and also presents videos that other viewers have given high ratings to. This allows users to select the content that best suits them from a vast amount of video data.
[0786] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0787] Step 1:
[0788] The user uploads video data to the system using their device. The uploaded video data is sent to the server. This data becomes the input for subsequent processing.
[0789] Step 2:
[0790] The server analyzes the received video data using TensorFlow. The analysis extracts visual features from the video and generates related symbolic information (text) based on these features. This process provides specific tag information for the video data.
[0791] Step 3:
[0792] The server automatically tags the extracted visual features and symbolic information. These tags are stored in MongoDB and output as a tagged database that will serve as the basis for future searches and recommendations.
[0793] Step 4:
[0794] The user enters a search query from their device. The query is received as a prompt that can be processed using natural language processing technology (e.g., "I'm looking for recommended videos that contain both comedy and heartwarming scenes."). This prompt becomes the input to the server.
[0795] Step 5:
[0796] The server uses Elasticsearch to match tag information in the database with search queries and quickly retrieve relevant video data. This process outputs search results for the video that best suits the user's request.
[0797] Step 6:
[0798] The server generates personalized recommendation lists based on viewing history and search queries. To achieve this, it analyzes past viewing activity and selects highly relevant content. This process results in an optimized video recommendation list for the user.
[0799] Step 7:
[0800] The generated search results and recommendation lists are displayed in a device application built with React Native. Users can then select and view video content that interests them based on the information presented.
[0801] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0802] Embodiments of the present invention include a system incorporating a user terminal, a central processing server, a network connecting them, and an emotion engine that recognizes the user's emotional state. This system allows the user to efficiently search for and manage video data corresponding to their emotions.
[0803] When a user selects video data through their device and uploads it to the system, this data is sent to a server. The server analyzes the received data using an AI model and extracts visual features and text information. Based on this information, automatic tagging is performed and stored in a database. This process enables efficient searching of video content.
[0804] Furthermore, the system incorporates an emotion engine that analyzes the user's emotions in real time while they are viewing videos. The emotion engine reads the user's emotional state from their facial expressions and voice and sends this information to the server. The server can then use this emotion information to recommend video content that is appropriate for the user's emotions and to adjust the search results.
[0805] As a concrete example, when a user views "soothing animal videos," the system's emotion engine analyzes the user's facial expressions and reactions. If the user appears relaxed and happy, the server prioritizes displaying similar relaxing video content. This process allows users to easily find video content that suits their emotional state.
[0806] Furthermore, by incorporating analysis information from the emotion engine into the tagging process, the tagging accuracy of video data is improved, resulting in more personalized search results. This makes it possible to provide users with the information they are looking for more quickly and accurately.
[0807] In this way, this system enables flexible management and retrieval of video data that takes into account the user's emotional state, thereby improving the user experience. By having the entire system work together to provide information that meets the user's needs, it can significantly improve efficiency in work and personal activities.
[0808] The following describes the processing flow.
[0809] Step 1:
[0810] The user uses a terminal to select video data and initiate the upload to the system. The terminal prepares the video data and transfers it to the server via the network.
[0811] Step 2:
[0812] The server processes the video data received from the terminal, using an AI model to extract visual features from the data. These extracted features are then used for tagging.
[0813] Step 3:
[0814] The server converts the audio information from the video data into text and analyzes the content within the video. This analysis is performed to generate highly relevant tags.
[0815] Step 4:
[0816] The server integrates visual features and text information and performs automatic tagging. This process assigns keywords related to the video data and stores them in a database.
[0817] Step 5:
[0818] When a user operates the device and begins watching a video, the emotion engine activates and starts analyzing the user's facial expressions and vocal characteristics.
[0819] Step 6:
[0820] The emotion engine analyzes the user's emotional state in real time and sends the obtained emotional information to the server.
[0821] Step 7:
[0822] The server selects video content that matches the user's emotions based on the received emotional information. In this process, videos with tags that match the user's emotions are given priority.
[0823] Step 8:
[0824] The user enters a search query on their device to search for a specific video. The device then sends this query to the server.
[0825] Step 9:
[0826] The server integrates search queries, tagging information, and user sentiment information to retrieve the most suitable video data from the database.
[0827] Step 10:
[0828] The search results are listed, ranked according to the user's emotional state, and sent to the device.
[0829] Step 11:
[0830] The terminal displays ranked search results received from the server to the user, presenting them in a visually easy-to-understand format.
[0831] Step 12:
[0832] If a user needs a summary of a specific video, they can request the server to generate the summary through their device.
[0833] Step 13:
[0834] The server uses generation technology to create a summary of the selected video content and provides it to the terminal.
[0835] Step 14:
[0836] Users can view summaries generated on their devices and quickly grasp the necessary information.
[0837] These steps enable users to efficiently search for and watch videos that resonate with their emotions, and to quickly understand the content they need.
[0838] (Example 2)
[0839] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0840] In today's digital society, it is crucial for users to quickly and accurately find appropriate digital content that matches their individual emotional state. However, conventional systems lacked the means to analyze user emotions in real time and recommend appropriate content, hindering improvements in the user experience. Furthermore, systems capable of dynamic content recommendation based on emotions were limited, and there was a need for information provision tailored to the individual needs of users.
[0841] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0842] In this invention, the server includes means for providing a terminal for users to upload data, means for using machine learning techniques to analyze the received data and extract its characteristics, means for automatically associating and storing identification information with the data, means for analyzing the user's emotional state, means for selecting and providing relevant information based on the emotional state, and means for dynamically adjusting the search results of the data according to the user's emotions. This enables advanced data recommendation and search that takes into account the user's emotional state, resulting in a significant improvement in the user experience.
[0843] "A terminal for users to upload data" refers to an electronic device used to transmit digital data specified by the user to the system.
[0844] "Means of using machine learning technology" refers to technologies that apply artificial intelligence algorithms to analyze received digital data and automatically identify and extract visual and textual information.
[0845] "Means for automatically associating and storing identification information with data" refers to the process of automatically assigning tags and metadata to analyzed digital data and saving that data to a storage device such as a database.
[0846] "Means for analyzing a user's emotional state" refers to technologies that analyze data such as a user's facial expressions and voice to identify the emotions they are expressing.
[0847] "Means of selecting and providing relevant information based on the aforementioned emotional state" refers to a process that automatically selects and presents digital content appropriate to the user's emotions based on the analyzed emotions of that user.
[0848] "Means of dynamically adjusting search results in response to user sentiment" refers to a mechanism that takes user sentiment data into consideration, modifies normal search results in real time, and prioritizes displaying the information most relevant to the user.
[0849] To implement this invention, a system is required in which a user, a terminal, and a server work together. The user can access the system using the terminal and select video data that they wish to view or upload. The terminal has an intuitive interface, making it easy for the user to operate.
[0850] In analyzing video data, the server utilizes machine learning techniques to analyze the received video data. Specifically, it can use widely used deep learning frameworks such as TensorFlow or PyTorch. The server uses these to extract visual features and text information from the video and automatically assigns identification information as tags. This automated tagging process allows the data to be efficiently stored in a database, enabling rapid searching and retrieval later on.
[0851] Furthermore, the server is equipped with an emotion engine to analyze the user's emotional state. This engine analyzes facial expressions and voice data transmitted from the user's device, allowing it to understand the user's emotional state in real time. Facial recognition technology and voice analysis technology are used for the analysis.
[0852] Based on the acquired emotional information, the server recommends video content that is appropriate for the user's current mood. In this recommendation process, the content recommendation algorithm reflects the emotional data and prioritizes providing the most relevant information to the user. For example, if a user is relaxing after watching "soothing animal videos," the server will continue to prioritize displaying videos with similar relaxing effects.
[0853] As a concrete example, here is an example of a prompt statement that utilizes a generative AI model:
[0854] "Please search for relaxing animal videos. If the user smiles, please recommend similar videos."
[0855] This significantly improves the user experience and enables content viewing tailored to individual emotional states. The integrated operation of the entire system makes managing and searching for digital content even more efficient.
[0856] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0857] Step 1:
[0858] The user selects digital data and uploads it to the system from the terminal. The input is a video file selected by the user, which is sent to the server via the terminal's interface. The output is the data arriving at the server. The terminal provides the user with a simple and intuitive UI, making file selection and uploading easy.
[0859] Step 2:
[0860] The server analyzes the received video data. The input is a video file, which is then fed into an AI model to extract visual features and textual information. Deep learning frameworks such as TensorFlow and PyTorch are used for data processing, and automatically tagged data is generated as output. The server efficiently stores this information in a database.
[0861] Step 3:
[0862] The system receives data to analyze the user's emotional state in real time. Input includes facial expressions and voice data sent from the device, which the emotion engine processes to determine the user's emotional state. Emotional information is generated as output and sent to the server. Facial recognition and voice tone analysis technologies are applied in the analysis.
[0863] Step 4:
[0864] The server uses user sentiment information to recommend content. It uses sentiment information and video content from its database as input. A content recommendation algorithm selects videos that match the sentiment, generating a list of candidates to display to the user. The server then sends this list to the user's device.
[0865] Step 5:
[0866] The user views recommended video content on their device. The input is a list of content sent from the server, from which the user can select and begin playback. The output further collects the user's emotional state and sends it back to the system as feedback to help improve future recommendations. This allows for continuous improvement of the user experience.
[0867] (Application Example 2)
[0868] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0869] Modern video streaming services are required to appropriately select and efficiently deliver the video content that users want, but content recommendations that take into account the user's emotional state are not being adequately implemented. As a result, user satisfaction is low, and the value of video streaming services is not being fully realized.
[0870] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0871] In this invention, the server includes means for analyzing the user's facial expression and voice information to determine their emotional state, means for recommending video content based on the determined emotional state, and means for providing search results to the user's terminal. This makes it possible to recommend video content based on the user's emotional state, enabling the efficient and effective provision of the video content the user desires.
[0872] "Video data" refers to digital data containing visual information, such as videos and images.
[0873] "Means of receiving" refers to mechanisms and methods for taking in data from an external source, such as acquiring information via networks or communication devices.
[0874] "Methods for analysis and extraction of visual and textual features" refers to methods for analyzing the content of video data and extracting features from image and textual information.
[0875] "Methods for automatic tagging" refer to the process of automatically assigning identifiable words or labels to data based on extracted features.
[0876] "Means of storing information in a database" refers to a system for long-term storage of structured information.
[0877] "A means of receiving search queries and searching for relevant video data" refers to a method of taking in information requests from users and finding corresponding video data.
[0878] "Means of providing to the user's terminal" refers to functions for displaying or delivering search results and recommended content to the user's device.
[0879] "A means of analyzing a user's facial expression and voice information to determine their emotional state" refers to a process of identifying a user's emotions at a given time by analyzing their facial movements and tone of voice.
[0880] "Methods for recommending video content based on emotional state" refer to technologies that take into account the determined emotions and select and present appropriate videos.
[0881] This system is implemented through a network that includes the user's smart device and a server. Video data viewed by the user is transmitted to the server in real time via a device such as a smartphone. The server analyzes the received video data and extracts visual features and text information. Feature extraction is performed using deep learning libraries such as OpenCV and TensorFlow.
[0882] Based on the analyzed features, a generative AI model automatically tags the data and stores it in a structured database. Furthermore, software equipped with an emotion engine detects changes in the user's facial expressions and voice to determine their emotional state. This analysis utilizes the smartphone's hardware, including its camera and microphone.
[0883] Once the user's emotional state is determined, the system uses that information to recommend video content that best suits the user's current emotions. This makes it easy for users to find content that resonates with their feelings.
[0884] For example, if the system detects that a user is crying while watching an emotionally moving film, it will recommend another film that is both touching and heartwarming. If the system detects laughter in the user's voice, it will then display a more entertaining comedy.
[0885] Here are some examples of prompts generated using a generative AI model. Instructions such as "If the user smiles, please display the following recommended content" or "If the user shows an emotion such as crying, please recommend a heartwarming movie" are used.
[0886] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0887] Step 1:
[0888] The user selects the video content they want to watch via their smart device. The device receives the video data and sends it to the server in real time. In this step, the video file is the input, and the output is data sent to the server in streaming format.
[0889] Step 2:
[0890] The server uses an AI model to extract visual features and text information to analyze the received video data. The input is video data, and the output is a list of features. Image analysis and natural language processing are performed using libraries such as OpenCV and TensorFlow, and the features are extracted as basic information for automatic tagging.
[0891] Step 3:
[0892] The server automatically tags the extracted visual and text features using a generative AI model. This tagging information is stored in a database. The input is a list of features, and the output is a data structure with automatic tags. The database storage process takes place at this stage.
[0893] Step 4:
[0894] When a user views video using their device, the device uses its camera and microphone to collect facial and audio information in real time. The collected data is sent to a server. The input consists of facial and audio data, and the output is the transmission of a data stream to the server.
[0895] Step 5:
[0896] The server analyzes the received facial expression and voice information to determine the user's emotional state. The input consists of facial expression data and voice data, and the output is a determination result indicating the emotional state (e.g., joy, sadness, etc.). The emotion engine performs this data calculation.
[0897] Step 6:
[0898] The server recommends video content that matches the user's current mood based on the determined emotional state. The input is the emotional state determination result and video database information, and the output is a list of recommended video content. A generative AI model is used for this recommendation process.
[0899] Step 7:
[0900] The terminal provides the user with video content recommended by the server. The user selects the content they want to watch next from the recommendation list displayed on the terminal. The input is the recommendation list received from the server, and the output is the video content selected by the user.
[0901] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0902] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0903] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0904] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0905] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0906] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0907] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0908] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0909] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0910] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0911] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0912] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0913] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0914] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0915] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0916] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0917] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0918] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0919] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0920] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0921] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[0922] The following is further disclosed regarding the embodiments described above.
[0923] (Claim 1)
[0924] A means of receiving video data,
[0925] A means of analyzing received video data and extracting visual and textual features,
[0926] A means for performing automatic tagging based on the characteristics,
[0927] A means of saving automatically tagged data to a database,
[0928] A means of receiving search queries and searching for related video data,
[0929] A means of providing search results to the user's terminal,
[0930] A system that includes this.
[0931] (Claim 2)
[0932] The system according to claim 1, which uses multiple modal learning techniques to integrate visual features and text information.
[0933] (Claim 3)
[0934] The system according to claim 1, which generates a summary of selected video content using a generation technique for providing the generated summary to a user terminal.
[0935] "Example 1"
[0936] (Claim 1)
[0937] A device for receiving video data,
[0938] A device for analyzing received video data and extracting visual attributes and text information,
[0939] A device for automatically tagging based on the said attribute,
[0940] A device for storing automatically tagged data on an information recording medium,
[0941] A device for receiving a search request and searching for video data related to the request,
[0942] A device for providing search results to a user-operated device,
[0943] A device for generating a video summary based on analyzed information and providing the summary to a user-operated device,
[0944] A system that includes this.
[0945] (Claim 2)
[0946] The system according to claim 1, which uses multiple modal learning techniques to integrate visual attributes and text information.
[0947] (Claim 3)
[0948] The system according to claim 1, which generates a summary of selected video material using a generation technique for providing the generated summary to a user operating device.
[0949] "Application Example 1"
[0950] (Claim 1)
[0951] A device that receives video data,
[0952] A device that analyzes received video data and extracts visual features and symbolic information,
[0953] A device that automatically tags based on the extracted information,
[0954] A device that automatically tags data and saves it to a storage device,
[0955] A device that receives search input and searches for related video data,
[0956] A device that provides search results to the user's terminal,
[0957] A device that generates a list of recommended video content,
[0958] A device that provides personalized recommendations using viewing history,
[0959] An information processing system that includes this.
[0960] (Claim 2)
[0961] An information processing system according to claim 1, which uses multiple learning techniques to integrate visual features and symbolic information.
[0962] (Claim 3)
[0963] The information processing system according to claim 1, which generates a summary of selected video content using a generation technology for providing the generated summary to a user terminal.
[0964] "Example 2 of combining an emotion engine"
[0965] (Claim 1)
[0966] A means of providing a terminal for users to upload data,
[0967] A means of using machine learning techniques to analyze received data and extract characteristics,
[0968] A means for automatically associating and storing identification information with data,
[0969] A means of analyzing the user's emotional state,
[0970] A means of selecting and providing relevant information based on the aforementioned emotional state,
[0971] A means of dynamically adjusting search results based on user sentiment,
[0972] A system that includes this.
[0973] (Claim 2)
[0974] The system according to claim 1, which uses multidimensional learning techniques to integrate visual characteristics and textual information.
[0975] (Claim 3)
[0976] The system according to claim 1, which creates a summary of selected information using generation technology and provides it to a terminal.
[0977] "Application example 2 of combining emotional engines"
[0978] (Claim 1)
[0979] A means of receiving video data,
[0980] A means of analyzing received video data and extracting visual and textual features,
[0981] A means for performing automatic tagging based on the characteristics,
[0982] A means of saving automatically tagged data to a database,
[0983] A means of receiving search queries and searching for related video data,
[0984] A means of providing search results to the user's terminal,
[0985] A means for analyzing the user's facial expressions and voice information to determine their emotional state,
[0986] A means of recommending video content based on the determined emotional state,
[0987] A system that includes this.
[0988] (Claim 2)
[0989] The system according to claim 1, which uses multiple modal learning techniques to integrate visual features and text information, and further comprises recommendation techniques based on emotional states.
[0990] (Claim 3)
[0991] The system according to claim 1, which provides a generated summary to a user terminal and uses generation technology to suggest video content that matches the user's emotional state. [Explanation of Symbols]
[0992] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of receiving video data, A means of analyzing received video data and extracting visual and textual features, A means for performing automatic tagging based on the characteristics, A means of saving automatically tagged data to a database, A means of receiving search queries and searching for related video data, A means of providing search results to the user's terminal, A system that includes this.
2. The system according to claim 1, which uses multiple modal learning techniques to integrate visual features and text information.
3. The system according to claim 1, which generates a summary of selected video content using a generation technique for providing the generated summary to a user terminal.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A