AI-powered real-time content metadata creation

AGI and machine learning models generate metadata for audio-video content in real-time, addressing the inefficiencies and spoiler risks of traditional methods by limiting training data to the current playback position, ensuring accurate and spoiler-free responses.

JP2026516687APending Publication Date: 2026-05-26SONY GROUP CORP

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SONY GROUP CORP
Filing Date
2024-02-06
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing audio-video content metadata creation methods are time-consuming and require significant effort, and chatbots trained on entire programs may inadvertently spoil the ending of ongoing content.

Method used

Utilizing artificial general intelligence (AGI) and machine learning models like ChatGPT4 to create metadata dynamically, limiting training data to the current playback position to avoid spoilers, and generating responses based on relevant document portions.

Benefits of technology

Enables efficient, spoiler-free metadata creation for audio-video content, providing accurate and timely responses to viewer queries without revealing future content details.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026516687000001_ABST
    Figure 2026516687000001_ABST
Patent Text Reader

Abstract

The system receives user queries seeking information related to audio-video (AV) programs (302), which are processed by machine learning (ML) models such as artificial intelligence (AL) chatbots (304). The chatbot returns a conversational response to the query (306). To prevent the chatbot from spoiling the ending of an AV program for the user midway through, the chatbot is supplied with only the corresponding text portion up to the point the user is watching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to non-routine electronic glossary solutions that are inherently rooted in computer technology and bring specific technical improvements, specifically related to the creation of real-time content metadata using AI.

Background Art

[0002] Audio-video (AV) content features many real and / or imaginary characters and concepts. Some people may have difficulty remembering all the details of a long AV program.

[0003] Searches can be performed to fill in the forgotten details. However, Internet searches require entering effective search terms, and furthermore, in the "hit or miss" process where the user has to examine multiple links in detail, which may or may not contain accurate or complete information, relevant information may not be obtained.

Summary of the Invention

[0004] This principle understands that when responding to viewer queries, it is usually necessary for someone to create the metadata of AV content so that it can be used. This process is very time-consuming and requires a great deal of effort to predict questions and possible responses. Furthermore, the response needs to be output in an audible form, and text-to-speech conversion may be required to read the response aloud to the viewer. This principle further understands that this problem is exacerbated in less popular works.

[0005] As will be further understood herein, the above concerns can be addressed by using artificial general intelligence (AGI), such as machine learning (ML) models like chatbots, to automatically create metadata, retrieve it from a large corpus of documents, and train this metadata to process viewer (user) queries about AV content. However, since chatbots are typically trained on documents covering entire programs, this principle recognizes that the chatbot may inadvertently spoil the ending of an AV program that a viewer is watching.

[0006] Therefore, from one perspective, this principle relates to a backend server that interacts with content players such as Netflix, Prime Video, and Audible running on client devices. The server receives requests for information related to content such as movies, shows, TV series, ebooks, or audiobooks. The server uses chatbot technology such as ChatGPT4, which uses a Large-Scale Language Model (LLM)-trained AGI to scan available character information and create responses.

[0007] For example, a user might ask a question about a character in AV content while watching it. The content player recognizes this as a query and relays it to the server along with the current playback position within the content. The server may have access to an overall summary of the content, possibly a rough synapsis, but not to detailed summaries of individual parts of the content. Therefore, it may be necessary to "load" all available print information about the content into an ML model or a GAI model (such as a chatbot) to create a more detailed response to the query. However, to avoid providing a response that reveals the content to the user, the relevant documents loaded into the AI ​​are limited to the point in the content corresponding to the user's current playback position, or the AI ​​is programmed to ignore loaded AV content information related to parts of the AV content after the user's current playback position when creating the response. The response is then provided to the user's content player to answer the query.

[0008] In one embodiment, the device includes at least one computer memory containing instructions, which are not transient signals, and the instructions are executable by at least one processor to present audio-video (AV) content on a display. The instructions are executable to receive at least one query related to the AV content and send this query, along with instructions for the current location within the AV content, to at least one server. The servers run at least one machine learning (ML) model. The instructions are executable to receive a response to the query from the ML model.

[0009] In an example embodiment, the device may also include a server that can include at least one processor configured to receive queries and instructions for the current location, run an ML model, and generate a response that includes information about AV content up to the current location.

[0010] In some implementations, the server's processor can be configured to search a corpus of documents for information related to AV content in response to queries. The processor can collect only the document portions in the corpus related to the AV content up to the current location and train an ML model on the document portions in the corpus related to the AV content up to the current location. The processor can be configured to run the ML model and generate a response.

[0011] In other implementations, the server's processor can be configured to search a corpus of documents for information related to AV content in response to queries, and to collect documents within the corpus that are relevant to the AV content. The processor can be configured to train an ML model on the documents within the corpus that are relevant to the AV content. In these implementations, the processor can be configured to run the ML model and generate a response using only the document portion up to the current location. For this purpose, the server's processor can be configured to train an ML model using only the document portion up to the current location.

[0012] In further implementations, the server's processor can be configured to search a corpus of documents for information relevant to AV content before a query is made, and to collect documents within the corpus that are relevant to the AV content. The server's processor can be configured to train an ML model on documents within the corpus that are relevant to the AV content. The processor can run the ML model to generate response sets for each of several query candidates. Each response set contains multiple responses, each of which corresponds to a location within the AV content, where each location is a position in a set of terms and / or sounds and / or images within the AV content. The server's processor can receive instructions for a query and the current location, select the response from each response set for the query candidate that best matches the query and the current location, and return each of these responses to the user system.

[0013] An ML model may include at least one generative pre-training transformer.

[0014] In another embodiment, the method includes presenting audio-video (AV) content on a display. The method includes receiving a query related to the AV content, transmitting the query and the current location within the AV content to at least one server on a wide area network, and presenting the server's response to the query.

[0015] In another embodiment, a device such as a server includes at least one processor configured to receive at least one query related to audio-video (AV) content and an indication of the current location within the AV content from at least one user system. The processor is configured to perform machine learning (modeling) to generate a response containing information about the AV content up to the current location only, and to return the response to the user system.

[0016] Details of this application, both in terms of its structure and operation, can be best understood by referring to the attached drawings, which indicate similar elements with similar reference numerals. [Brief explanation of the drawing]

[0017] [Figure 1] This is a block diagram of an example system based on this principle. [Figure 2] This figure shows an example of a user system including an AV player that works in conjunction with one or more servers. [Figure 3] This figure shows an example of AV player logic in flowchart format based on this principle. [Figure 4] This figure shows an example of the first server logic in flowchart format based on this principle. [Figure 5] This figure shows a second example of server logic in flowchart format based on this principle. [Figure 6] This figure shows examples of screenshots that the user system can present in an example embodiment. [Figure 7] This figure shows an example of the first user system logic in the flowchart format shown in Figure 6. [Figure 8] This figure shows a second example of user system logic in the flowchart format shown in Figure 6. [Figure 9] Figure 6 shows an example of server logic in a flowchart format. [Figure 10] This figure shows another example of server logic in a flowchart format based on this principle. [Modes for carrying out the invention]

[0018] This disclosure relates to a computer ecosystem including a computer network, which may generally include consumer electronic (CE) devices. The systems described herein may include server and client components connected via a network to exchange data with each other. The client components may include one or more computer devices, including portable computers such as portable televisions (e.g., smart TVs, internet-enabled TVs), laptop computers, and tablet computers, as well as other mobile devices, including smart headphones, smartphones, and further examples described later. These client devices may operate in a variety of operating environments. For example, some client computers may use, as an example, an operating system from Microsoft, or a Unix operating system, or an operating system from Apple or Google. Using these operating environments, one or more browsing programs may be run, such as browsers, which may access websites hosted by Internet servers described later, created by Microsoft, Google, Mozilla, or other browser programs.

[0019] The server and / or gateway may include one or more processors that execute instructions to configure the server to send and receive data over a network such as the Internet. Alternatively, the client and server may be connected via a local intranet or a virtual private network. The server or controller may also be exemplified by a game console such as Sony PlayStation®, a personal computer, etc.

[0020] Between a client and a server, information can be exchanged via a network. For this purpose and for security, the server and / or the client can include a firewall, a load balancer, temporary storage, and a proxy, as well as other network infrastructure to enhance reliability and security.

[0021] As used herein, an instruction means a computer-executable step for processing information within a system. An instruction can be implemented in software, firmware, or hardware, and can include any type of program step performed by a component of the system.

[0022] A processor can be a general-purpose single-chip or multi-chip processor that can execute logic using various lines such as address lines, data lines, and control lines, as well as registers and shift registers.

[0023] The software modules described by the flowchart, and the user interfaces herein, can include various subroutines, procedures, etc. Without limiting the present disclosure, the logic disclosed as being performed by a particular module can be redistributed to other software modules and / or combined into a single module and / or made available within a shareable library.

[0024] The principles described herein can be implemented as hardware, software, firmware, or a combination thereof, and thus the exemplary components, blocks, modules, circuits, and steps are described in terms of their functional aspects.

[0025] In addition to those suggested above, the logic blocks, modules, and circuits described below may be implemented, or performed, by, a combination of, a general-purpose processor, a digital signal processor (DSP), a field-programmable gate array (FPGA), or other programmable logic devices such as application-specific integrated circuits (ASICs), discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. The processor may be implemented by a combination of a controller, a state machine, or a computer device.

[0026] The functions and methods described below, when implemented in software, can be written in a suitable language such as C# or C++, but are not limited to such languages, and can be stored in or transmitted through computer-readable storage media such as random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), compact disk read-only memory (CD-ROM), or other optical disk storage such as digital purpose discs (DVDs), magnetic disk storage, or other magnetic storage devices including removable thumb drives. A certain connection can constitute computer-readable media. Such connections may include, as an example, wired cables including optical fibers, coaxial cables, digital subscriber lines (DSL), and twisted pair cables.

[0027] Components included in one embodiment may also be used in any suitable combination in other embodiments. For example, any of the various components described and / or shown in the figures herein may be combined, replaced, or excluded from other embodiments.

[0028] A "system having at least one of A, B, and C (similarly, a "system having at least one of A, B, or C", and a "system having at least one of A, B, and C")" includes systems having only A, only B, only C, both A and B, both A and C, both B and C, and / or all of A, B, and C.

[0029] Referring specifically to Figure 1, an example ecosystem 10 is shown which may include one or more of the device examples described above and further described below, based on the present principle. The first device example included in system 10 is a consumer electronic (CE) device configured as a primary display device example, which in the illustrated embodiment is, but is not limited to, an audio-video display device (AVDD) 12 such as an internet-enabled TV having a TV tuner (corresponding to a set-top box that controls the TV). AVDD 12 may be an Android®-based system. Alternatively, AVDD 12 may be, for example, a computerized internet-enabled ("smart") phone, a tablet computer, a notebook computer, a computerized internet-enabled watch or other computerized wearable device, a computerized internet-enabled bracelet or other computerized internet-enabled device, a computerized internet-enabled music player, a computerized internet-enabled headphones or other computerized internet-enabled implantable device such as an implantable skin-worn device. In any case, it should be understood that AVDD12 and / or other computers described herein are configured to implement the principle (for example, by communicating with other CE devices to implement the principle, by executing the logic described herein, and by performing any other functions and / or operations described herein).

[0030] Therefore, AVDD12 can be constructed using some or all of the components shown in Figure 1 to implement such principles. For example, AVDD12 may include one or more displays 14 that are touch-enabled or not, which can receive user input signals via touch to the display, and which can be implemented by high-definition or ultra-high-definition "4K" or higher-definition flat screens. AVDD12 may also include one or more speakers 16 that output audio according to these principles, and at least one further input device 18, such as an audio receiver / microphone that inputs audible commands to AVDD12 to control AVDD12. AVDD example 12 may further include one or more network interfaces 20 for communicating over at least one network 22 such as the Internet, WAN, LAN, PAN, etc., under the control of one or more processors 24. Therefore, the interface 20 may be a Wi-Fi transceiver, which is an example of a wireless computer network interface such as a mesh network transceiver, but is not limited to these. Interface 20 may be, but is not limited to, a Bluetooth transceiver, a Zigbee transceiver, an IrDA transceiver, a wireless USB transceiver, a wired USB, a wired LAN, a power line, or MoCA. It should be understood that the processor 24 controls AVDD 12 to implement this principle, including other elements of AVDD 12 described herein, such as controlling the display 14 to present an image and receiving input from the display 14. Furthermore, network interface 20 may also be other suitable interfaces, such as a wired or wireless modem or router, or a 5G-based wireless telephone transceiver, or a Wi-Fi transceiver as described above.

[0031] In addition to the above, AVDD12 may also include one or more input ports 26, such as a high-definition multimedia interface (HDMI) port or a USB port, for physically connecting to another CE device (for example, using a wired connection), and / or a headphone port for connecting headphones to AVDD12 so that audio can be presented to the user through headphones from AVDD12. For example, the input ports 26 may be connected to a cable or satellite source 26a of audio video content via wired or wireless connection. Thus, source 26a may be, for example, a standalone or integrated set-top box or a satellite receiver. Alternatively, source 26a may be a game console or a disc player.

[0032] AVDD12 may further include one or more computer memories 28, such as non-transient signal disk storage or solid storage, which may be embodied within the chassis of AVDD, either as a standalone device, or as a personal video recording device (PVR) or video disc player for playing AV programs inside or outside the chassis of AVDD, or as a removable storage medium. In some embodiments, AVDD12 may also include a location or place receiver, such as a cell phone receiver, GPS receiver and / or altimeter 30, which is configured to receive geographic location information from, for example, at least one satellite or cell phone tower and provide it to the processor 24, and / or is configured to cooperate with the processor 24 to determine the altitude at which AVDD12 is located. However, it should be understood that the location of AVDD12 may also be determined in all three dimensions, for example, using another suitable location receiver other than a cell phone receiver, GPS receiver and / or altimeter, in accordance with this principle.

[0033] Continuing the description of AVDD12, in some embodiments, AVDD12 may also include one or more cameras 32, which may be digital cameras such as thermal cameras, webcams, and / or cameras integrated with AVDD12 and controllable by a processor 24 to collect photographs / images and / or videos in accordance with this principle. AVDD12 may also include a Bluetooth transceiver 34 and other NFC elements 36 that communicate with other devices using Bluetooth and / or Near Field Communication (NFC) technology, respectively. An example of an NFC element is a radio frequency identification (RFID) element.

[0034] Furthermore, AVDD12 may also include one or more auxiliary sensors 37 that provide input to the processor 24 (e.g., motion sensors such as accelerometers, gyroscopes, cyclometers, or magnetic sensors; infrared (IR) sensors for receiving IR commands from remote control devices; optical sensors; speed and / or gait sensors; gesture sensors (for sensing gesture commands, etc.)). AVDD12 may also include a wireless TV broadcast port 38 for receiving OTH TV broadcasts that provide input to the processor 24. In addition to the above, AVDD12 may also include an infrared (IR) transmitter and / or IR receiver and / or IR transceiver 42, such as an IR Data Communications Association (IRDA) device. A battery (not shown) may also be provided to power AVDD12.

[0035] Furthermore, in some embodiments, AVDD12 may include a graphics processing unit (GPU) and / or a field-programmable gate array (FPGA) 39. AVDD12 can utilize the GPU and / or FPGA 39 for artificial intelligence processing, such as training a neural network and performing the operation of the neural network (e.g., inference) in accordance with the present principle. However, if the processor 24 is a central processing unit (CPU), the processor 24 can also be used for artificial intelligence processing.

[0036] Continuing to refer to Figure 1, System 10 may include, in addition to AVDD 12, one or more other computer device types that may include some or all of the components of AVDD 12 shown. In one example, a first device 44 and a second device 46 are shown, which may include some or all of the components of AVDD 12. Fewer or more devices may be used than those shown.

[0037] In the illustrated example, to illustrate this principle, we assume that all three devices 12, 44, and 46 are members of a local network within a residence 48, for example, indicated by a dashed line.

[0038] An unspecified first example of a device 44 may include one or more touch-sensitive surfaces 50, such as a touch-enabled video display, that receive user input signals via touch on the display. The first device 44 may also include one or more speakers 52 that output audio according to this principle, and at least one further input device 54, such as an audio receiver / microphone that inputs audible commands to the first device 44 to control the device 44. The first example of a device 44 may also include one or more network interfaces 56 for communication over a network 22 under the control of one or more processors 58. Thus, the interfaces 56 may be Wi-Fi transceivers, including mesh network interfaces, which are examples of wireless computer network interfaces, but are not limited to these. The processors 58 should be understood to control the first device 44 to implement this principle, including other elements of the first device 44 described herein, such as controlling the display 50 to present images and receiving input from the display 50. Furthermore, the network interface 56 can also be, for example, a wired or wireless modem or router, or a wireless telephone transceiver, or another suitable interface such as the Wi-Fi transceiver mentioned above.

[0039] In addition to the above, the first device 44 may also include one or more input ports 60, such as an HDMI port or a USB port, for physically connecting to another computer device (for example, using a wired connection), and / or a headphone port for connecting headphones to the first device 44 so that audio can be presented to the user through headphones from the first device 44. The first device 44 may further include one or more tangible computer-readable storage media 62, such as disk storage or solid storage. Also, in some embodiments, the first device 44 may also include a location or place receiver, such as a cell phone receiver and / or a GPS receiver and / or altimeter 64, which is configured to receive geographic location information using triangulation from, for example, at least one satellite and / or cell phone tower and provide it to the device processor 58, and / or is configured to determine the altitude at which the device 44 is located in cooperation with the device processor 58. However, it should be understood that the location of the first device can also be determined, for example, in all three dimensions, using another suitable location receiver other than a cell phone receiver and / or a GPS receiver and / or altimeter in accordance with this principle.

[0040] Continuing the description of the first device 44, in some embodiments, the first device 44 may include one or more cameras 66, which may be, for example, infrared cameras, webcams, or other digital cameras. The first device 44 may also include a Bluetooth transceiver 68 and other NFC elements 70 that communicate with other devices using Bluetooth and / or Near Field Communication (NFC) technology, respectively. An example of an NFC element is a radio frequency identification (RFID) element.

[0041] Furthermore, the first device 44 may also include one or more auxiliary sensors 72 that provide input to the CE device processor 58 (e.g., motion sensors such as accelerometers, gyroscopes, wheel rotation recorders, or magnetic sensors, infrared (IR) sensors, optical sensors, speed and / or gait sensors, gesture sensors (for sensing gesture commands, etc.)). The first device 44 may also include one or more climate sensors 74 that provide input to the device processor 58 (e.g., barometers, humidity sensors, wind sensors, light sensors, temperature sensors, etc.) and / or one or more biosensors 76, and other sensors. In addition to the above, in some embodiments, the first device 44 may also include an infrared (IR) transmitter and / or IR receiver and / or IR transceiver 42, such as an IR Data Communications Association (IRDA) device. A battery may also be provided to power the first device 44. The device 44 can communicate with AVDD 12 through any of the communication modes and associated components described above.

[0042] The second device 46 may include some or all of the components described above.

[0043] Next, referring to the at least one server 80 described above, this server 80 includes at least one server processor 82, at least one computer memory 84 such as disk storage or solid storage, and at least one network interface 86 which enables communication with other devices in Figure 1 via the network 22 under the control of the server processor 82, and in practice facilitates communication between the server, controller and client devices in accordance with this principle. Note that the network interface 86 may also be other suitable interfaces such as a wired or wireless modem or router, a Wi-Fi transceiver, or a wireless telephone transceiver.

[0044] Therefore, in some embodiments, the server 80 can be an internet server, and in an example embodiment, such a function can be performed by including a "cloud" function so that the devices of system 10 can access the "cloud" environment via the server 80. Alternatively, a game console or other computer in the same room as, or nearby than, the other devices shown in Figure 1, can implement the server 80.

[0045] The device described later may incorporate some or all of the elements mentioned above.

[0046] The methods described herein can be implemented as software instructions executed by a processor, a suitably configured application-specific integrated circuit (ASIC), or a field-programmable gate array (FPGA) module, or in any other convenient form understood by those skilled in the art. Software instructions can be embodied in non-transient devices such as CD-ROMs or flash drives, where applicable. Alternatively, software code instructions can be embodied in a transient configuration such as wireless or optical signals, or through downloads over the Internet.

[0047] This principle can employ a variety of machine learning models, including deep learning models. Machine learning models based on this principle can utilize various algorithms trained using methods including supervised learning, unsupervised learning, semi-supervised learning, reinforcement learning, feature learning, self-learning, and other forms of learning. Examples of such algorithms that can be implemented by computer circuits include one or more neural networks, such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), and RNNs of the type known as long short-term memory (LSTM) networks. Support vector machines (SVMs) and Bayesian networks can also be considered examples of machine learning models.

[0048] Therefore, as understood herein, performing machine learning can involve training a model on training data after accessing the training data so that the model can process further data and perform inference. Thus, an artificial neural network / artificial intelligence model trained through machine learning may include an input layer, an output layer, and multiple hidden layers between these, configured and weighted to perform inference about appropriate outputs.

[0049] In specific embodiments, the ML model employs a transformer-based neural network architecture, such as a generative pre-trained transformer trained on a large dataset of text, to generate human-like text that can be converted to speech in response to queries.

[0050] Figure 2 shows a user system 200 that can implement any of the above-described devices for displaying AV content supplied by the AV player 202 on a video display 204 and / or speaker 206 under the control of one or more processors 208 that access instructions on a computer storage medium.

[0051] AV content can include, for example, movies, TV programs, music (audio only), and text such as e-book texts.

[0052] In short, the processor 208 can run one or more machine learning (ML) models 210. The processor 208 can receive queries about AV content via one or more input devices 212 such as a microphone, keyboard, or keypad, and send these queries via one or more network interfaces 214 to one or more Internet or other wide area network services 216 having one or more processors 218 running one or more ML models 220, and return the responses to the queries to the user system 200 via one or more network interfaces 222. The server processor 218 can access a document corpus 224, such as network sites on the Internet, when running the ML models. In some examples, transformer-based neural network architectures, such as generative pre-trained transformers that establish a "chatbot," can implement the server's ML models 220.

[0053] Figure 3 shows the logic that the user system can execute by first training the ML model 210 of the user system in Figure 2 to recognize queries. This training can be done by training the model on a set of training phrases for which the ground truth indicates whether each phrase is a query or not.

[0054] After training, the process proceeds to block 302, where input is received in the user system and sent to the ML model 210 to determine whether it is a query. If the input is recognized as a query about AV content, which is typically presented to the user system 200, in block 304, this query is sent to the server 216 along with the identity of the AV content and the user's current location within the AV content. In block 306, a response is received, and in block 308, this response is presented audibly and / or visually and / or tactilely on the user system.

[0055] Next, refer to Figure 4, which shows a first example of how server 216 responds to a query. In block 400, it enters a do loop for a specific AV content, and in block 402, it receives a query and its current location from the user system (including the identification of the relevant AV content).

[0056] Moving to block 404, the server searches corpus 224, shown in Figure 2, for all documents related to the AV content. However, in the example in Figure 4, moving to block 406, only the document portions related to the AV content up to the current location are collected, and the ML model 220 on server 216 is trained only on these collected portions. Once trained, the ML model 220 generates conversational responses to queries and returns the responses to the user system in block 408.

[0057] Figure 5 shows another way to respond to a query. In block 500, a do loop is entered for a specific AV content, and in block 502, the query and current location are received from the user system.

[0058] Moving to block 504, the server searches corpus 224, shown in Figure 2, for all documents related to the AV content. In the example in Figure 5, the entire document is collected by the server in block 506, but only the document portion related to the AV content up to the current location is used to train the ML model 220 on server 216. Once trained, the ML model 220 generates conversational responses to queries and returns the responses to the user system in block 508.

[0059] Figure 6 shows a further feature in which a user interface (UI) 600 can be presented on the user system display 204, prompting the user at 602 whether they want a query response that includes spoiler information, i.e., information related to AV program segments after the current position. A selector 604 can be provided that allows the user to choose whether or not to include spoiler information.

[0060] Figure 7 shows the user system logic of the UI 600 in Figure 6 that can be executed by the AV player 202 shown in Figure 2. If it is determined in state 700 that spoilers are desired, the received query is sent in block 702 to server 216 along with commands for the server to train the chatbot about all available information related to the AV content. If spoilers are not desired, the logic calls Figure 3 in state 704.

[0061] Figure 8 shows another way the user system handles spoilers. If it is determined in state 800 that spoilers are desired, the received query is sent in block 802 to server 216 to train the chatbot on all available information related to the AV content, but this does not include instructions on the current location. If spoilers are not desired, the logic calls Figure 3 in state 804.

[0062] Figure 9 shows the server-side logic according to Figures 7 to 9. In state 900, if a command is received to train the chatbot on all available documents related to AV content associated with the query, or equivalently if the current location is not received along with the query, the server trains the chatbot on all available content in block 902 and returns a response to the query. On the other hand, if state 900 indicates that the response should not contain spoilers, the server can execute the logic in Figure 4 or Figure 5 in state 904.

[0063] Figure 10 shows another pre-query logic that the server can execute to train the chatbot before a query is received. Before the query, once AV content is selected in block 1000, block 1002 searches the corpus for documents related to the AV content from block 1000. In block 1004, the server's ML model 220 is trained on these documents.

[0064] Proceed to block 1006 and run ML model 220 to generate a response set for each of the multiple query candidates. Each response set for each query candidate contains multiple responses, each of which corresponds to a specific location within the AV content. Note that each location should be understood as a position within a set of terms and / or sounds and / or images within the AV content.

[0065] In block 1008, after training, a query is received along with an indication of the current location. The process proceeds to block 1010, where the system selects the appropriate response from the respective response sets for the closest query that matches the user's query and is closest to the current location indicated by the user system.

[0066] Regarding training corpora, audiobooks can use text such as the text of printed books. Closed captions for videos and video scripts can also be used. Furthermore, dialogue recognized from audio related to books or films can be used. The tone of the dialogue can undergo sentiment analysis to collect details from spoken language. Thus, training can be multimodal, incorporating not only text but also audio-text and video (where available).

[0067] Training and / or query responses can follow listeners or viewers in real time as they consume content, and a large-scale language model can be used to create a corpus of material to investigate. Thus, when closed captions are unavailable, speech-text can be used for any dialogue and (where available) audio descriptions. The text can also be used to find entry points in the printed text to determine the playback position in the printed content.

[0068] ML models can be run on the user's local computer, on a mobile phone, or as part of a player application.

[0069] While this principle has been described with reference to several embodiments, these embodiments are not intended to be limiting, and it will be understood that the subject matter claimed herein can also be implemented using a variety of other configurations. [Explanation of symbols]

[0070] 400 AV Contents DO 402 query, received user's current location within content. Search the corpus for all documents related to 404 AV content. 406 Collect only the document portion up to the current location for training. 408 Response returned for query

Claims

1. It is a device, The computer has at least one computer memory containing instructions, which are not temporary signals, and the instructions are: Audio-video (AV) content is displayed on the screen, Receiving at least one query related to the aforementioned AV content, The query is sent to at least one server running at least one machine learning (ML) model, along with an instruction for the current location within the AV content. The response to the aforementioned query is received from the ML model. Thus, it is executable by at least one processor. A device characterized by the following features.

2. The system comprises at least one server, and the at least one server is Upon receiving the aforementioned query and the current location instruction, The ML model is executed to generate a response that includes information about the AV content only up to the current location. Including at least one processor configured as follows: The apparatus according to claim 1.

3. The processor of the server is In response to the aforementioned query, the document corpus is searched for information related to the AV content, Collect only the document portion within the corpus related to the AV content up to the current location, The ML model is trained on the document portion within the corpus related to the AV content up to the current location. The ML model is executed to generate the response. The apparatus according to claim 2, configured as follows.

4. The processor of the server is In response to the aforementioned query, the document corpus is searched for information related to the AV content, Collect documents from the corpus related to the aforementioned AV content, The ML model is trained on the documents in the corpus related to the AV content. The ML model is executed and a response is generated using only the document portion up to the current position. The apparatus according to claim 2, configured as follows.

5. The processor of the server is configured to train the ML model using only the document portion up to the current location. The apparatus according to claim 4.

6. The processor of the server is Prior to the aforementioned query, the document corpus is searched for information related to the AV content, Collect documents from the corpus related to the aforementioned AV content, The ML model is trained on the documents in the corpus related to the AV content. The ML model is executed to generate a set of responses for each of the multiple query candidates, each response set containing multiple responses, each of which corresponds to a position within the AV content, and each position is a position in a sequence of terms and / or sounds and / or images within the AV content. Upon receiving the aforementioned query and the current location instruction, Select the respective response from the set of responses for the aforementioned query and the query candidate that most closely matches the current position. The apparatus according to claim 2, configured as follows.

7. The ML model includes at least one generative pre-trained transformer. The apparatus according to claim 1.

8. Displaying audio-video (AV) content on a display, Receiving queries related to the aforementioned AV content, The query and the current location within the AV content are transmitted to at least one server on the wide area network. To present the response from the server to the aforementioned query, A method characterized by including the following.

9. The server receives the query and the current location information, The server executes a machine learning (ML) model to generate a response that includes information about the AV content up to the current location, To return the response to the device associated with the display, The method according to claim 8, including the method described in claim 8.

10. In response to the aforementioned query, the document corpus is searched for information related to the AV content, To collect only the document portion within the corpus related to the AV content up to the current location, Training the ML model with respect to the document portion within the corpus related to the AV content up to the current location, Executing the aforementioned ML model to generate the aforementioned response, The method according to claim 9, including the method described in claim 9.

11. In response to the aforementioned query, the document corpus is searched for information related to the AV content, Collecting documents within the corpus related to the aforementioned AV content, Training the ML model with respect to the documents in the corpus related to the AV content, The ML model is executed to generate a response using only the document portion up to the current position, The method according to claim 9, including the method described in claim 9.

12. This includes training the ML model using only the document portion up to the current location. The method according to claim 11.

13. Prior to the aforementioned query, the document corpus is searched for information related to the AV content, Collecting documents within the corpus related to the aforementioned AV content, Training the ML model with respect to the documents in the corpus related to the AV content, The ML model is executed to generate a set of responses for each of the multiple query candidates, each set of responses comprising multiple responses, each of which corresponds to a position within the AV content, and each position being a position in a sequence of terms and / or sounds and / or images within the AV content. Receiving the aforementioned query and the instruction for the current location, Selecting each response from the respective sets of responses for the aforementioned query and the query candidate that most closely matches the current position, To return each of the above responses, The method according to claim 9, including the method described in claim 9.

14. The ML model includes at least one generative pre-trained transformer. The method according to claim 9.

15. It is a device, The system receives at least one query related to audio-video (AV) content and an indication of the current location within the AV content from at least one user system. Machine learning (model) is executed to generate a response that includes information about the AV content up to the current location, The response is returned to the user system. A system comprising at least one processor configured as follows: A device characterized by the following features.

16. The aforementioned processor, In response to the aforementioned query, the document corpus is searched for information related to the AV content, Collect only the document portion within the corpus related to the AV content up to the current location, The ML model is trained on the document portion within the corpus related to the AV content up to the current location. The ML model is executed to generate the response. The apparatus according to claim 15, configured as follows.

17. The aforementioned processor, In response to the aforementioned query, the document corpus is searched for information related to the AV content, Collect documents from the corpus related to the aforementioned AV content, The ML model is trained on the documents in the corpus related to the AV content. The ML model is executed and a response is generated using only the document portion up to the current position. The apparatus according to claim 15, configured as follows.

18. The processor is configured to train the ML model using only the document portion up to the current position. The apparatus according to claim 17.

19. The aforementioned processor, Prior to the aforementioned query, the document corpus is searched for information related to the AV content, Collect documents from the corpus related to the aforementioned AV content, The ML model is trained on the documents in the corpus related to the AV content. The ML model is executed to generate a set of responses for each of the multiple query candidates, each response set containing multiple responses, each of which corresponds to a position within the AV content, and each position is a position in a sequence of terms and / or sounds and / or images within the AV content. Upon receiving the aforementioned query and the current location instruction, Select the respective response from the set of responses for the aforementioned query and the query candidate that most closely matches the current position. The apparatus according to claim 15, configured as follows.

20. The ML model includes at least one generative pre-trained transformer. The apparatus according to claim 15.