Streaming media system and streaming media server
By adaptively generating predicted frames and encoding streaming data using an interactive frame prediction model in the streaming media system, the problem of difficulty in quickly providing high-resolution frames in game streaming services is solved, and efficient user experience improvement is achieved.
Patent Information
- Application Number
- CN202011445628.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-04-16
- Filing Date
- 2020-12-08
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2040-12-08
AI Technical Summary
The prior art is difficult to quickly provide high-resolution frames in game streaming services, resulting in poor user experience.
By using an interactive frame prediction model in a streaming media system, predicted frames are adaptively generated based on user input and selectively encoded streaming media data using these predicted frames, increasing the compression rate.
It realizes sending high-resolution frames to users in real time or near real time, improving the user experience of game streaming services.
Smart Images

Figure CN113542749B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of Korean Patent Application No. 10-2020-0045787 filed on April 16, 2020 in the Korean Intellectual Property Office (KIPO), the entire contents of which are incorporated herein by reference. Technical Field
[0003] Example embodiments relate generally to providing streaming data, and more particularly to a streaming system and a streaming server that provide streaming services with an interactive scenario. Background Art
[0004] Interactive streaming services such as game streaming are one of the next generation fields that have emerged recently. Recently, high-spec games have been introduced, and game streaming services have become important because client devices cannot run high-spec games. In game streaming services, it is important to provide users with high-resolution frames quickly. Summary of the invention
[0005] Some example embodiments provide a streaming media system capable of adaptively using an interactive frame prediction model based on user input.
[0006] Some example embodiments provide a method of providing an interactive streaming media service, the method being capable of adaptively using an interactive frame prediction model based on user input.
[0007] According to some example embodiments, a streaming media system includes a streaming media server and a client device. The streaming media server is configured to train an interactive frame prediction model based on streaming media data, user input, and metadata associated with the user input; encode the streaming media data by selectively using a prediction frame generated based on the trained interactive frame prediction model; and send the trained interactive frame prediction model and the encoded streaming media data. The client device is configured to receive the trained interactive frame prediction model and the encoded streaming media data; and decode the encoded streaming media data based on the trained interactive frame prediction model to provide the restored streaming media data to a user.
[0008] According to some example embodiments, a streaming media system includes a streaming media server and a client device. The streaming media server is configured to: select a target interactive frame prediction model from among a plurality of interactive frame prediction models; train the target interactive frame prediction model based on streaming media data, user input, and metadata associated with the user input; generate a prediction frame by using the target interactive frame prediction model; encode the streaming media data by selectively using the prediction frame; and send the encoded streaming media data. The client device is configured to: receive the plurality of interactive frame prediction models and the encoded streaming media data; select the target interactive frame prediction model from among the plurality of interactive frame prediction models; and decode the encoded streaming media data based on the target interactive frame prediction model to provide the restored streaming media data to a user.
[0009] According to some example embodiments, in a method for providing an interactive streaming service, an interactive frame prediction model trained by a streaming server based on streaming data, user input, and metadata associated with the user input, and encoded streaming data are provided by the streaming server to a client device, the client device decodes the encoded streaming data based on the trained interactive frame prediction model received from the streaming server to generate restored streaming data, and the client device displays the restored streaming data to a user via the display.
[0010] Therefore, the streaming media system and related methods can increase or improve the compression rate by generating predicted frames based on the use of an interactive frame prediction model and encoding the streaming media data based on the predicted frames. Therefore, the streaming media system and related methods can send high-resolution frames to users in real time or near real time. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Illustrative, non-limiting example embodiments will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings.
[0012] Figure 1 An electronic device is shown in a network environment according to an example embodiment.
[0013] Figure 2 is a block diagram illustrating an electronic device according to an example embodiment.
[0014] Figure 3 is a flow chart illustrating a method of providing an interactive streaming service according to an example embodiment.
[0015] Figure 4 is a diagram showing a method according to an example embodiment Figure 3 Flowchart of the operation of the streaming media server in.
[0016] Figure 5 is a diagram showing a method according to an example embodiment Figure 3 Flowchart of the operation of a client device in FIG.
[0017] Figure 6 is a diagram showing a method according to an example embodiment Figure 1 A block diagram of an example of a streaming media server in FIG.
[0018] Figure 7 is a diagram showing a method according to an example embodiment Figure 1 A block diagram of an example of a client device in FIG.
[0019] Fig. 8A and Figure 8B is used to describe the Figure 6 An illustration of an example of a neural network in .
[0020] Fig. 9 is a diagram showing a method according to an example embodiment Figure 6 A block diagram of an example of an encoder in a streaming media server.
[0021] Fig.10 is a diagram showing a method according to an example embodiment Figure 7 A block diagram of an example of a decoder in a client device.
[0022] Fig.11 According to an example embodiment, Figure 6 The operation of the encoder and Figure 7 The operation of the decoder in .
[0023] Fig.12 is an example operation of a streaming media server according to an example embodiment.
[0024] Fig.13 are example operations of a client device according to an example embodiment.
[0025] Fig.14 is a diagram showing a method according to an example embodiment Figure 1 A block diagram of an example of a streaming media server in FIG.
[0026] Fig.15 A streaming media system according to an example embodiment is shown.
[0027] Fig.16 is a diagram showing a method according to an example embodiment Fig.15 A block diagram of an example of a streaming media card in FIG.
[0028] Fig.17 is a flow chart illustrating the operation of a client device according to an example embodiment.
[0029] Fig.18A and Fig.18B Example operations of a client device are respectively shown.
[0030] Fig.19 Example operations of a client device according to example embodiments are shown.
[0031] Fig. 20 Example operations of a client device according to example embodiments are shown.
[0032] Fig.21 A training operation of a streaming media server according to an example embodiment is shown.
[0033] Fig. 22 is a block diagram illustrating an electronic system according to example embodiments. DETAILED DESCRIPTION
[0034] Example embodiments will be described more fully hereinafter with reference to the accompanying drawings.
[0035] Figure 1 An electronic device is shown in a network environment according to an example embodiment.
[0036] Reference Figure 1 , the electronic device 101 in the network environment 100 may include a bus 110, a processor 120, a memory 130, an input / output (I / O) interface 150, a display 160, and a communication interface 170. According to some example embodiments, the electronic device 101 may omit at least one of the aforementioned elements, or may further include other elements. The bus 110 may include a circuit for connecting, for example, elements 110 to 170 and transmitting communications (e.g., control messages or data) between elements 110 to 170. The processor 120 may include one or more of a central processing unit (CPU), an application processor (AP), and a communication processor (CP). The processor 120 performs operations or data processing for control and / or communication of at least one other element of, for example, the electronic device 101.
[0037] The processor 120 and / or any portion thereof (e.g., a processing unit), and alternatively other computer devices (e.g., servers and streaming media cards), may be implemented by one or more of the following instances: a processing circuit such as hardware including logic circuits; a hardware / software combination such as a processor running software as described in the above embodiments; or a combination thereof.
[0038] The memory 130 may include a volatile memory and / or a non-volatile memory. The memory 130 may store, for example, instructions or data associated with at least one other element of the electronic device 101. According to some example embodiments, the memory 130 may store software and / or a program 140. The program 140 may include, for example, at least one of a kernel 141, middleware 143, an application programming interface (API) 145, and / or an application program (or "application") 147, etc. At least some of the kernel 141, middleware 143, and API 145 may be referred to as an operating system (OS). The kernel 141 may control or manage, for example, system resources (e.g., bus 110, processor 120, memory 130, etc.) for executing operations or functions implemented in other programs (e.g., middleware 143, API 145, or application program 147).
[0039] The kernel 141 provides an interface through which the middleware 143, the API 145, and / or the application program 147 access different components of the electronic device 101 to control or manage system resources.
[0040] The middleware 143 may be used as a medium to allow, for example, the API 145 or the application 147 to exchange data with the kernel 141. In addition, the middleware 143 may process one or more task requests received from the application 147 based on a priority. For example, the middleware 143 may give priority to at least one application 147 in using system resources (e.g., the bus 110, the processor 120, the memory 130, etc.) of the electronic device 101, and may process one or more task requests.
[0041] The API 145 is an interface used by the application 147 to control functions provided by the kernel 141 or the middleware 143, and may include, for example, at least one interface or function (e.g., instruction) for file control, window control, image processing, or character control. The I / O interface 150 may, for example, transfer instructions or data input from a user or another external device to other components of the electronic device 101, or output instructions or data received from other components of the electronic device 101 to the user or another external device.
[0042] The display 160 may include, for example, a liquid crystal display (LCD), a light emitting diode (LED) display, an organic light emitting diode (OLED) display, a micro-electromechanical system (MEMS) display, or an electronic paper display. The display 160 may, for example, display various contents (e.g., text, images, videos, icons, and / or symbols, etc.) to the user. The display 160 may include a touch screen and receive, for example, a touch, gesture, proximity, or hovering input by using an electronic pen or a part of the user's body.
[0043] The communication interface 170 establishes communication between the electronic device 101 and an external device (e.g., the first external electronic device 102, the second external electronic device 104, or the server 106). For example, the communication interface 170 may be connected to the network 162 through wireless communication or wired communication to communicate with the external device (e.g., the second external electronic device 104 or the server 106).
[0044] The wireless communication may include a cellular communication protocol using, for example, at least one of the following: Long Term Evolution (LTE), Advanced LTE (LTE-A), Code Division Multiple Access (CDMA), Wideband CDMA (WCDMA), Universal Mobile Telecommunications System (UMTS), Wireless Broadband (WiBro), Global System for Mobile Communications (GSM), etc. According to some example embodiments, the wireless communication may include at least one of Wireless Fidelity (WiFi), Bluetooth, Bluetooth Low Energy (BLE), Zigbee, Near Field Communication (NFC), Magnetic Secure Transmission (MST), Radio Frequency (RF), and Body Area Network (BAN). According to some example embodiments, the wireless communication may include a Global Navigation Satellite System (GNSS). The GNSS may include, for example, at least one of a Global Positioning System (GPS), a Global Navigation Satellite System (Glonass), a Beidou Navigation Satellite System (“Beidou”), and Galileo (European Global Satellite-based Navigation System). In the following, “GPS” may be used interchangeably with “GNSS”. Wired communication may include, for example, at least one of Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), Recommended Standard 232 (RS-232), power line communication, Plain Old Telephone Service (POTS), and the like.
[0045] The network 162 may include a telecommunication network, for example, at least one of a computer network (eg, a local area network (LAN) or a wide area network (WAN)), the Internet, and a telephone network.
[0046] The first external electronic device 102 and the second external electronic device 104 may each be a device of the same type or a different type as the electronic device 101 .
[0047] According to some example embodiments, some or all of the operations performed by the electronic device 101 may be performed in another electronic device or multiple electronic devices (e.g., electronic device 102 or 104 or server 106). According to some example embodiments of the present invention, when the electronic device 101 must automatically or upon request perform a function or service, the electronic device 101 may request another device (e.g., electronic device 102 or 104 or server 106) to perform at least some functions associated with the function or service without performing the function or service, or to perform at least some functions associated with the function or service in addition to performing the function or service. Another electronic device (e.g., electronic device 102 or 104 or server 106) may perform the requested function or additional function and pass the execution result to the electronic device 101. Then, the electronic device 101 may process or further process the received result to provide the requested function or service. To this end, for example, cloud computing, distributed computing, or client-server computing technology may be used.
[0048] exist Figure 1 In the embodiment, electronic devices 101, 102 and 104 may be referred to as client devices, and server 106 may be referred to as a streaming media server.
[0049] Figure 2 is a block diagram illustrating an electronic device according to an example embodiment.
[0050] Reference Figure 2 , the electronic device 201 may be formed Figure 1 The entire electronic device 101 or Figure 1 A portion of an electronic device 101 is shown.
[0051] The electronic device 201 may include one or more processors (e.g., an application processor (AP)) 210, a communication module 220, a user identification module (SIM) 224, a memory 230, a sensor module 240, an input device 250, a display 260, an interface 270, an audio module 280, a camera module 291, a power management module 295, a battery 296, an indicator 297, and a motor 298.
[0052] The processor 210 controls a plurality of hardware or software components connected to the processor 210 by driving an operating system (OS) or an application program, and performs processing and operations on various data. The processor 210 may be implemented, for example, with a system on a chip (SoC). According to some example embodiments of the inventive concept, the processor 210 may include a graphics processing unit (GPU) and / or an image signal processor. The processor 210 may include Figure 2At least some of the elements shown in (e.g., cellular module 221). The processor 210 loads instructions or data received from at least one of the other elements (e.g., non-volatile memory) into the volatile memory to process the instructions or data, and stores the result data in the non-volatile memory.
[0053] The communication module 220 may have the same or similar configuration as the communication interface 170. The communication module 220 may include, for example, a cellular module 221, a WiFi module 223, a Bluetooth (BT) module 225, a GNSS module 227, an NFC module 228, and a radio frequency (RF) module 229.
[0054] The cellular module 221 may provide, for example, a voice call, a video call, a text service, or an Internet service through a communication network. According to some example embodiments, the cellular module 221 identifies and authenticates the electronic device 201 in the communication network by using a SIM 224 (e.g., a SIM card). According to some example embodiments, the cellular module 221 may perform at least one of the functions that may be provided by the processor 210.
[0055] According to some example embodiments, the cellular module 221 may include a communication processor (CP). According to some example embodiments, at least some (e.g., two or more) of the cellular module 221, the WiFi module 223, the BT module 225, the GNSS module 227, and the NFC module 228 may be included in one integrated chip (IC) or an IC package.
[0056] The RF module 229 may, for example, send and receive communication signals (e.g., RF signals). The RF module 229 may include a transceiver, a power amplifier module (PAM), a frequency filter, a low noise amplifier (LNA), or an antenna. According to some example embodiments, at least one of the cellular module 221, the WiFi module 223, the BT module 225, the GNSS module 227, and the NFC module 228 may send and receive RF signals through a separate RF module.
[0057] SIM 224 may include, for example, a card (including a SIM or an embedded SIM), and may include unique identification information (eg, an integrated circuit card identifier (ICCID) or subscriber information (eg, an International Mobile Subscriber Identity (IMSI)).
[0058] The memory 230 (eg, the memory 130 ) may include, for example, an internal memory 232 and / or an external memory 234 .
[0059] The internal memory 232 may include, for example, at least one of a volatile memory (e.g., dynamic random access memory (DRAM), static RAM (SRAM), synchronous dynamic RAM (SDRAM), etc.), a non-volatile memory (e.g., one-time programmable read-only memory (OTPROM), programmable ROM (PROM), erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), etc.), a mask ROM, a flash ROM, a flash memory, and a solid-state drive (SSD).
[0060] The external memory 234 may also include a flash disk, such as a compact flash (CF), secure digital (SD), micro SD, ultra micro SD, extreme digital (xD), multimedia card (MMC) or memory stick. The external memory 234 may be functionally or physically connected to the electronic device 201 through various interfaces.
[0061] The sensor module 240 measures a physical quantity or senses an operating state of the electronic device 201 to convert the measured or sensed information into an electric signal.
[0062] The sensor module 240 may, for example, include at least one of a posture sensor 240A, a gyroscope sensor 240B, a pressure sensor 240C, a magnetic sensor 240D, an acceleration sensor 240E, a grip sensor 240F, a proximity sensor 240G, a color sensor 240H (e.g., a red / green / blue (RGB) sensor), a biometric sensor 240I, a temperature / humidity sensor 240J, an illuminance sensor 240K, and an ultraviolet (UV) sensor 240M.
[0063] Additionally or alternatively, the sensor module 240 may include an electronic nose sensor (not shown), an electromyogram (EMG) sensor (not shown), an electroencephalogram (EEG) sensor (not shown), an electrocardiogram (ECG) sensor (not shown), an infrared (IR) sensor, an iris sensor, and / or a fingerprint sensor. The sensor module 240 may also include a control circuit for controlling at least one sensor included therein. In some example embodiments, the electronic device 201 may further include a processor configured to control the sensor module 240 as part of the processor 210 or separately from the processor 210 to control the sensor module 240 during a dormant state of the processor 210.
[0064] The input device 250 may include, for example, a touch panel 252, a (digital) pen sensor 254, a key 256, or an ultrasonic input device 258, but example embodiments are not limited thereto. The input device 250 may be configured to receive commands from outside the electronic device 201. The touch panel 252 may use at least one of a capacitive type, a resistive type, an IR type, or an ultrasonic type. The touch panel 252 may further include a control circuit.
[0065] The touch panel 252 may also include a tactile layer to provide a tactile response to the user. The (digital) pen sensor 254 may include an identification sheet as part of the touch panel 252 or a separate identification sheet. The key 256 may also include a physical button, an optical key, or a keypad. The ultrasonic input device 258 senses the ultrasonic wave generated by the input device through a microphone (e.g., microphone 288) and checks the data corresponding to the sensed ultrasonic wave.
[0066] Display 260 (e.g., display 160) may include panel 262, hologram device 264, projector 266, and / or a control circuit for controlling them. Panel 262 may be implemented as flexible, transparent, or wearable. Panel 262 may be configured in a module with touch panel 252.
[0067] According to some example embodiments, the panel 262 may include a pressure sensor (or a "force sensor" used interchangeably hereinafter) capable of measuring the intensity of pressure by the user's touch. The pressure sensor may be implemented integrally with the touch panel 252, or may be implemented as one or more sensors separated from the touch panel 252. The hologram device 264 may display a stereoscopic image in the air by utilizing interference of light. The projector 266 may display an image on a screen by projection of light. The screen may be located inside or outside the electronic device 201.
[0068] The interface 270 may include an HDMI 272, a USB 274, an optical interface 276, or a D-SUB 278. The interface 270 may be included in Figure 1 The communication interface 170 is shown in FIG. Additionally or alternatively, the interface 270 may include a mobile high-definition link (MHL) interface, a SD / multimedia card (MMC) interface, or an infrared data association (IrDA) interface.
[0069] The audio module 280 can bidirectionally convert sound and electrical signals. At least one element of the audio module 280 may include Figure 1 The input / output interface 150 is shown. The audio module 280 can process sound information input or output through the speaker 282, the receiver 284, the earphone 286 or the microphone 288.
[0070] The camera module 291 is, for example, a device capable of capturing still images or moving images, and according to some example embodiments, the camera module 291 may include one or more image sensors (e.g., a front sensor or a rear sensor), a lens, an image signal processor (ISP), or a flash (e.g., an LED, a xenon lamp, etc.).
[0071] The power management module 295 manages power of the electronic device 201. According to some example embodiments, the power management module 295 may include a power management integrated circuit (PMIC), a charger IC, or a battery gauge.
[0072] The PMIC may have a wired and / or wireless charging scheme. The wireless charging scheme may include a magnetic resonance type, a magnetic induction type, and an electromagnetic type, and may also include additional circuits for wireless charging, such as a coil loop, a resonant circuit, or a rectifier. The battery gauge may measure the remaining capacity of the battery 296 or measure the voltage, current, or temperature of the battery 296 during charging. The battery 296 may include, for example, a rechargeable battery and / or a solar cell.
[0073] Indicator 297 displays a specific state of electronic device 201 or a portion thereof (e.g., processor 210), such as a startup state, a message state, or a charging state. Motor 298 may convert an electrical signal into a mechanical vibration or generate a vibration or tactile effect. Electronic device 201 may include a processor 298 for supporting a mobile TV (e.g., a GPU) to broadcast a video according to a digital multimedia broadcasting (DMB), a digital video broadcasting (DVB), or a mediaFlo TM Devices that use standards to process media data.
[0074] Figure 3 is a flow chart illustrating a method of providing an interactive streaming service according to an example embodiment.
[0075] In the following, Figure 1 In the embodiment, the electronic device 101 may correspond to a client device, and the server 106 may correspond to a streaming media server.
[0076] Reference Figure 1 and Figure 3 , the streaming media server 106 connected to the client device 101 through the network 162 can send the encoded streaming media data and the interactive frame prediction model trained based on the streaming media data, the user input and the metadata associated with the user input to the client device 101 (operation S100). The interactive frame prediction model can be implemented with a neural network and can be stored in a memory in the streaming media server 106.
[0077] The processor in the streaming media server 106 can apply the user input, the learning streaming media data and the metadata to the interactive frame prediction model to train the interactive frame prediction model. When the training of the interactive frame prediction model is completed, the interactive frame prediction model can receive the user input, the streaming media data and the metadata as input, and can provide a predicted frame of the object frame of the streaming media data as output. The encoder in the streaming media server 106 can encode the streaming media data by selectively using the predicted frame, and can provide the encoded streaming media data to the client device 101.
[0078] The client device 101 may decode the encoded streaming media data based on the trained interactive frame prediction model received from the streaming media server 106 to generate restored streaming media data (operation S300). The client device 101 may display (provide) the restored streaming media data to the user through a display (operation S500). The user may apply user input to the streaming media data through the I / O interface.
[0079] Figure 4 is a diagram showing a method according to an example embodiment Figure 3 Flowchart of the operation of the streaming media server in.
[0080] Reference Figure 4 In order to send the encoded streaming media data and the interactive frame prediction model (operation S100), the processor in the streaming media server 106 can apply the frames of the streaming media data, user input and metadata to the neural network to train the interactive frame prediction model (operation S110).
[0081] When the training of the interactive frame prediction model is completed, the processor in the streaming server 106 can input the previous frame, user input and metadata to the interactive frame prediction model to generate a prediction frame of the object frame of the streaming data (operation S130). The previous frame is the previous frame of the object frame.
[0082] The encoder in the streaming server 106 may encode the streaming data by referring to the higher similarity frame selected from the previous frame and the predicted frame to generate the encoded streaming data (operation S150). The higher similarity frame has a higher similarity with the object frame.
[0083] The streaming media server 106 may send the encoded streaming media data to the client device 101 (operation S170 ).
[0084] Figure 5 is a diagram showing a method according to an example embodiment Figure 3 Flowchart of the operation of a client device in FIG.
[0085] Reference Figure 5In order to generate the recovered streaming data (operation S300), the processor in the client device 101 can apply the encoded streaming data, user input, and metadata to the trained interactive frame prediction model received from the streaming server 106 to generate a predicted frame (operation S310).
[0086] The decoder in the client device 101 may decode the encoded streaming media data based on the prediction frame to generate the restored streaming media data (operation S330). The decoder in the client device 101 may decode the object frame by referring to a higher similarity frame selected from the prediction frame and the previous frame of the encoded streaming media data.
[0087] The processor in the client device 101 may display (provide) the restored streaming media data to the user through the display in the client device 101 (operation S350 ).
[0088] Figure 6 is a diagram showing a method according to an example embodiment Figure 1 A block diagram of an example of a streaming media server in FIG.
[0089] Reference Figure 6 , the streaming media server 106 may include a processor 420, a memory 430, an operation server 440, and a training server 480. The processor 420, the memory 430, the operation server 440, and the training server 480 may be operably coupled to each other via a bus 410.
[0090] The operation server 440 may include a GPU 445 and an encoder 450. The training server 480 may store a neural network 485 that implements an interactive frame prediction model IFPM.
[0091] The memory 430 may store instructions. The processor 420 may execute the instructions stored in the memory 430 to control the operation server 440 and the training server 480 to perform specific operations.
[0092] The GPU 445 may generate streaming data SRDT associated with the virtual reality game, and may provide the streaming data SRDT to the buffer (BUF) 435 and the encoder 450. The buffer 435 may store the streaming data SRDT by frame and may provide the frame of the streaming data SRDT to the training server 480.
[0093] The training server 480 may apply the user input UIN, the metadata MDT associated with the user input UIN, and the streaming media data SRDT to the interactive frame prediction model IFPM to train the interactive frame prediction model IFPM. For example, in a game, the metadata MDT may be information for understanding the context of the game from the user's perspective. The metadata MDT may include the played map information, character information, and weapon information.
[0094] When the difference between the compression rate of the output of the interactive frame prediction model IFPM in response to the user input UIN, the metadata MDT, and the streaming media data SRDT and the compression rate having the expected value is within the reference range (or less than the reference value), the processor 420 may determine that the training of the interactive frame prediction model IFPM is completed. That is, the processor 420 may determine that the training of the interactive frame prediction model IFPM is completed in response to the difference between the compression rate of the predicted frame and the compression rate of the expected frame associated with the predicted frame being within the reference range.
[0095] The training server 480 may pre-train the interactive frame prediction model IFPM, or may train the interactive frame prediction model IFPM when the training server 480 receives the user input UIN and the metadata MDT.
[0096] When the training of the interactive frame prediction model IFPM is completed, the processor 420 controls the training server 480 to send the trained interactive frame prediction model IFPM to the client device 101 .
[0097] The trained interactive frame prediction model IFPM may provide a predicted frame PFR to the encoder 450 in response to the user input UIN, metadata MDT and streaming media data SRDT as inputs.
[0098] The encoder 450 can encode the object frame of the streaming data SRDT by referring to a higher similarity frame selected from the prediction frame PFR and the previous frame of the streaming data SRDT that has a higher similarity to the object frame to generate the encoded streaming data ESRDT, and the encoded streaming data ESRDT can be sent to the client device 101.
[0099] When the encoder 450 performs inter prediction or intra prediction, the encoder 450 may encode the object frame of the streaming media data SRDT by referring to a higher similarity frame selected from the prediction frame PFR and a previous frame of the streaming media data SRDT.
[0100] The streaming media server 106 may send the trained interactive frame prediction model IFPM and the encoded streaming media data ESRDT to the client device 101 via the communication interface included therein.
[0101] Figure 7 is a diagram showing a method according to an example embodiment Figure 1 A block diagram of an example of a client device in FIG.
[0102] Reference Figure 7 , the client device 101 may include a processor 120, a memory 130, an I / O interface 150, a display 160, and a communication interface 170. The processor 120, the memory 130, the I / O interface 150, the display 160, and the communication interface 170 may be coupled to each other via a bus 110.
[0103] The memory 130 may store instructions. The processor 120 may execute the instructions stored in the memory 130 to control the I / O interface 150, the display 160, and the communication interface 170 to perform specific operations.
[0104] The I / O interface 150 may receive the user input UIN, and may provide the user input UIN and metadata MDT associated with the user input UIN to the communication interface 170 .
[0105] The communication interface 170 can send user input UIN and metadata MDT to the streaming server 106, can receive the trained interactive frame prediction model IFPM and encoded streaming data ESRDT from the streaming server 106, can store the trained interactive frame prediction model IFPM in the memory 130, and can provide the encoded streaming data ESRDT to the decoder 175 in the processor 120.
[0106] The processor 120 may apply the user input UIN, the metadata MDT, and the encoded streaming data ESRDT to generate a predicted frame about the object frame of the encoded streaming data ESRDT, and the decoder 175 in the processor 120 may decode the object frame by referencing a higher similarity frame selected from the previous frame of the encoded streaming data ESRDT and the predicted frame of the encoded streaming data ESRDT to restore the encoded streaming data ESRDT, thereby generating a restored streaming data RSRDT. The processor 120 may provide the restored streaming data RSRDT to the user by displaying the restored streaming data RSRDT in the display 160.
[0107] Referring to the restored streaming data RSRDT displayed in the display 160, the user may participate in the game implemented by the restored streaming data RSRDT by applying the user input UIN to the restored streaming data RSRDT.
[0108] Fig. 8A and Figure 8Bis used to describe the Figure 6 An illustration of an example of a neural network in .
[0109] Reference Fig. 8A A general neural network may include: an input layer IL, multiple hidden layers HL1, HL2, ..., HLn, and an output layer OL.
[0110] The input layer IL may include i input nodes x 1 、x 2 , …, x i , where i is a natural number. Input data (e.g., vector input data) IDAT of length i can be input to the input node x 1 、x 2 , …, x i , so that each element in the streaming data SRDT, user input UIN and metadata MDT is input to the input node x 1 、x 2 , …, x i The corresponding input node in .
[0111] The plurality of hidden layers HL1, HL2, ..., HLn may include n hidden layers, where n is a natural number, and may include a plurality of hidden nodes h 1 1 、h 1 2 、h 1 3 ,…,h 1 m 、h 2 1 、h 2 2 、h 2 3 ,…,h 2 m 、h n 1 、h n 2 、h n 3 ,…,h n m For example, the hidden layer HL1 may include m hidden nodes h 1 1 、h 1 2 、h 1 3 ,…,h 1 m , the hidden layer HL2 may include m hidden nodes h 2 1 、h2 2 、h 2 3 ,…,h 2 m , and the hidden layer HLn may include m hidden nodes h n 1 、h n 2 、h n 3 ,…,h n m , where m is a natural number.
[0112] The output layer OL may include j output nodes y 1 ,y 2 , …, y j , where j is a natural number. Output node y 1 ,y 2 , …, y j Each output node in may correspond to a corresponding one of the classes to be classified. The output layer OL may output an output value (e.g., a class score or a simple score) or a predicted frame PFR associated with the input data of each class. The output layer OL may be referred to as a fully connected layer and may indicate, for example, a probability that the predicted frame PFR corresponds to the desired frame.
[0113] Fig. 8A The structure of the neural network shown can be represented by information about branches (or connections) between nodes shown as lines and weight values (not shown) assigned to each branch. Nodes within a layer may not be connected to each other, but nodes in different layers may be fully or partially connected to each other.
[0114] Each node (for example, node h 1 1 ) can receive the previous node (for example, node x 1 ), may perform a computation, estimation, or operation on the received output, and may output the result of the computation, estimation, or operation as an output to the next node (e.g., node h 2 1 ). Each node can calculate the value to be output by applying a specific function (e.g., a nonlinear function) to the input.
[0115] Generally, the structure of the neural network is set in advance, and the weight values for the connection between nodes are appropriately set using data whose category the data belongs to as a known answer. The data with known answers is called "training data", and the process of determining the weight values is called "training". The neural network "learns" during the training process. A set of independently trainable structures and weight values is called a "model", and the process of predicting the category to which the input data belongs through the model with determined weight values and then outputting the predicted value is called the "testing" process.
[0116] Reference Figure 8B , showing in detail the Fig. 8A An example of an operation performed by a node ND included in a neural network.
[0117] When N input a 1 、a 2 、a 3 , …, a N When provided to node ND, node ND can take N input a 1 、a 2 、a 3 , …, a N Respectively with the corresponding N weights w 1 、w 2 、w 3 ,…,w N By multiplication, N values obtained by the multiplication may be summed, an offset "b" may be added to the summed value, and an output value (eg, "z") may be generated by applying the value to which the offset "b" is added to a specific function "σ".
[0118] when Fig. 8A One layer of the neural network shown in Figure 8B When the M nodes ND are shown in , the output value of this layer can be obtained by formula 1.
[0119] W*A=Z (Formula 1)
[0120] In Equation 1, “W” represents the weights of all connections included in this layer and can be implemented in the form of an M*N matrix. “A” represents the N inputs a received by this layer. 1 、a 2 、a 3 , …, a N , and can be implemented in N*1 matrix form. "Z" represents the M output z output from this layer 1 、z 2 、z 3 ,…,z M , and can be implemented in the form of an M*1 matrix.
[0121] Fig. 9 is a diagram showing a method according to an example embodiment Figure 6 A block diagram of an example of an encoder in a streaming media server.
[0122] Reference Fig. 9 , the encoder 450 may include a mode decision block (MD) 451 , a compression block 460 , an entropy encoder (EC) 467 , a reconstruction block 470 , and a storage block (STG) 477 .
[0123] The mode decision block 451 may generate a first prediction frame PRE based on the current frame Fn and the reference frame REF, and may generate encoding information INF including a prediction mode according to a prediction operation, a result of the prediction operation, a syntax element, a context value, etc. The mode decision block 451 may include a motion estimation unit (ME) 452, a motion compensation unit (MC) 453, and an intra prediction unit (INTP) 454. The intra prediction unit 454 may perform intra prediction. The motion estimation unit 452 and the motion compensation unit 453 may be referred to as an inter prediction unit that performs inter prediction.
[0124] The compression block 460 may encode the current frame Fn to generate an encoded frame EF. The compression block 460 may include a subtractor 461, a transform unit (T) 463, and a quantization unit (Q) 465. The subtractor 461 may subtract the first prediction frame PRE from the current frame Fn to generate a residual frame RES. The transform unit 463 and the quantization unit 465 may transform and quantize the residual frame RES to generate an encoded frame EF.
[0125] The reconstruction (restoration) block 470 may be used to generate a reconstructed frame Fn' by inversely decoding the coded frame EF. The reconstruction block 470 may include an inverse quantization unit (Q -1 )471, inverse transform unit (T -1 )473 and adder 475.
[0126] The inverse quantization unit 471 and the inverse transform unit 473 may inversely quantize and inversely transform the encoded frame EF to generate a residual frame RES'. The adder 475 may add the residual frame RES' to the first predicted frame PRE to generate a reconstructed frame Fn'.
[0127] The entropy encoder 467 may perform lossless encoding on the encoded frame EF and the encoding information INF to generate the encoded streaming media data ESRDT. The reconstructed frame Fn' may be stored in the storage block 477 and may be used as another reference frame for encoding other frames.
[0128] The storage block 477 may store the previous frame Fn-1 and the predicted frame PFR output from the interactive frame prediction model IFPM, and the motion estimation unit 452 may perform motion estimation by referring to a higher similarity frame selected from the previous frame Fn-1 and the predicted frame PFR having a higher similarity with the object (current) frame Fn. That is, the encoder 450 may encode the object frame Fn by using the higher similarity frame selected from the previous frame Fn-1 and the predicted frame PFR to provide the encoded streaming media data ESRDT to the client device 101.
[0129] Fig.10 is a diagram showing a method according to an example embodiment Figure 7 A block diagram of an example of a decoder in a client device.
[0130] Reference Fig.10 The decoder 175 may include an entropy decoder (ED) 176, a prediction block 180, a reconstruction block 185, and a storage block (STG) 190. The decoder 175 may generate the restored streaming media data RSRDT by inversely decoding the encoded streaming media data ESRDT encoded by the encoder 450.
[0131] The entropy decoder 176 may decode the encoded streaming media data ESRDT to generate an encoded frame EF and encoding information INF.
[0132] The prediction block 180 may generate a second prediction frame PRE' based on the reference frame REF and the encoding information INF. The prediction block 180 may include Fig. 9 The motion compensation unit 453 and the intra-frame prediction unit 454 in the image are basically the same as the motion compensation unit 181 and the intra-frame prediction unit 183.
[0133] The reconstruction block 185 may include an inverse quantization unit 186, an inverse transform unit 187, and an adder 188. The reconstruction block 185 and the storage block 190 may be respectively Fig. 9 The reconstruction block 470 and the storage block 477 in FIG. 4 are substantially the same. The reconstructed frame Fn′ may be stored in the storage block 190 and may be used as another reference frame, or may be provided to the display 160 as recovered streaming media data RSRDT.
[0134] The storage block 190 may store the predicted frame PFR' provided from the interactive frame prediction model IFPM, and the prediction block 180 may generate a second predicted frame PRE' by using a higher similarity frame having a higher similarity with the reconstructed frame Fn' selected from the previous frame of the reconstructed frame Fn' and the predicted frame PFR' as a reference frame REF.
[0135] Fig.11 According to an example embodiment, Figure 6The operation of the encoder and Figure 7 The operation of the decoder in .
[0136] In an example embodiment, the encoder 450 and the decoder 175 may respectively perform encoding and decoding on a specific form of a group of pictures (GOP) structure. The GOP may conform to the standard defined by the Moving Picture Experts Group (MPEG). According to the above standard, the GOP may have three frames. The GOP may have a combination of I frames, P frames and / or B frames. For example, the GOP may have a repetition of "IB...BPB...BP". As another example, the GOP may have a repetition of "IP...P". The three frames may be intra-coded frames (I frames), predicted frames (P frames) or bidirectional predicted frames (B frames). The I frame may be an independent frame. The P frame may be a frame related to an I frame or a P frame. The B frame may be a frame related to at least one of an I frame and a P frame. For example, a B frame may be generated based on a higher similarity frame selected from an I frame and a P frame. The B frame may have a higher compression rate than the P frame, and the P frame may have a higher compression rate than the I frame.
[0137] During the training phase TRP, a GOP may have a repetition of "IPPP", while during the inference phase IFP, a GOP may have a repetition of "I B'B'B'". Here, the B' frames correspond to the predicted frames provided by the trained interactive frame prediction model IFPM.
[0138] Reference Figure 6 , Figure 7 and Fig.11 , the CPU 440 in the streaming server 106 sequentially generates frames F1, F2, F3 and F4. The interactive frame prediction model IFPM generates prediction frames F1(P'), F2(P') and F3(P') of each of the frames F1, F2 and F3 based on the user input UIN and the metadata MDT, and provides the prediction frames F1(P'), F2(P') and F3(P') to the encoder 450.
[0139] The encoder 450 encodes the frame F1 to generate an encoded frame F1(I), and encodes the object frame F2 by referring to a higher similarity frame selected from the previous frame F1 and the predicted frame F1(P') having a higher similarity to the object frame F2 to generate an encoded frame F2(B'). The encoder 450 encodes the object frame F3 by referring to a higher similarity frame selected from the previous frame F2(B') and the predicted frame F2(P') having a higher similarity to the object frame F3 to generate an encoded frame F3(B'), and encodes the object frame F4 by referring to a higher similarity frame selected from the previous frame F3(B') and the predicted frame F3(P') having a higher similarity to the object frame F4 to generate an encoded frame F4(B').
[0140] The streaming server 106 provides the coded frames F1(I), F2(B'), F3(B'), and F4(B') to the client device 101. The trained interactive frame prediction model IFPM in the client device 101 generates prediction frames F1(P'), F2(P'), and F3(P') for each of the coded frames F1(I), F2(B'), and F3(B') based on the user input UIN and the metadata MDT, and provides the prediction frames F1(P'), F2(P'), and F3(P') to the decoder 175.
[0141] The decoder 175 decodes the encoded frame F1(I) to provide the restored frame F1 to the display 160, and decodes the object frame F2(B') by referring to the higher similarity frame selected from the previous frame F1(I) and the predicted frame F1(P') having a higher similarity to the object frame F2(B') to provide the restored frame F2 to the display 160. The decoder 175 decodes the object frame F3(B') by referring to the higher similarity frame selected from the previous frame F2(B') and the predicted frame F2(P') having a higher similarity to the object frame F3(B') to provide the restored frame F3 to the display 160, and decodes the object frame F4(B') by referring to the higher similarity frame selected from the previous frame F3(B') and the predicted frame F3(P') having a higher similarity to the object frame F4(B') to provide the restored frame F4 to the display 160.
[0142] exist Fig.11 In the embodiment, when it is assumed that the performance of the interactive frame prediction model IFPM is 100%, the streaming server 106 sends the I frame of each GOP to the client device 101, and the client device 101 generates a prediction frame by using the interactive frame prediction model IFPM to provide the restored streaming data to the user.
[0143] Fig.12 is an example operation of a streaming media server according to an example embodiment.
[0144] Reference Fig.12 , the streaming server 106 may send multiple interactive frame prediction models IFPM1 (310), IFPM2 (320), and IFPM3 (330) to multiple client devices 101 and 301. The multiple interactive frame prediction models 310, 320, and 330 are associated with multiple domains in a game implemented by the streaming data.
[0145] The streaming server 106 may train a plurality of interactive frame prediction models 310 , 320 , and 330 , and may send the interactive frame prediction models 310 , 320 , and 330 to the client devices 101 and 301 after completing the training of the interactive frame prediction models 310 , 320 , and 330 .
[0146] The interactive frame prediction model 310 may be associated with a first domain in a game implemented by streaming data, the interactive frame prediction model 320 may be associated with a second domain in a game implemented by streaming data, and the interactive frame prediction model 330 may be associated with a third domain in a game implemented by streaming data.
[0147] Both client devices 101 and 301 can store interactive frame prediction models 310, 320 and 330 in the memory therein, and client device 101 can select interactive frame prediction model 310 among interactive frame prediction models 310, 320 and 330 to use the selected interactive frame prediction model 310 for decoding, while client device 301 can select interactive frame prediction model 320 among interactive frame prediction models 310, 320 and 330 to use the selected interactive frame prediction model 320 for decoding.
[0148] Fig.13 are example operations of a client device according to an example embodiment.
[0149] Reference Fig.13 , the client device 101 can obtain target streaming media data corresponding to the target resolution of the original streaming media data for the target domain.
[0150] In some example embodiments, the target domain may include a plurality of subdomains corresponding to at least one of the plurality of designated areas. The target domain may include a first subdomain 510a corresponding to a first area (e.g., a general venue area) and a second subdomain 510b corresponding to a second area (e.g., a dungeon A area).
[0151] The client device 101 may select a target interactive frame prediction model corresponding to a target resolution of the original streaming media data for a target domain. The target interactive frame prediction model may include a plurality of sub-interactive frame prediction models corresponding to a plurality of sub-domains. For example, the target interactive frame prediction model may include a first sub-interactive frame prediction model SUB_IFPM1 511 corresponding to a first sub-domain 510a and a second sub-interactive frame prediction model SUB_IFPM2 512 corresponding to a second sub-domain 510b.
[0152] The client device 101 may select the first sub-interactive frame prediction model 511 based on obtaining target streaming media data corresponding to a target resolution of the original streaming media data for a target domain.
[0153] Fig.14 is a diagram showing a method according to an example embodiment Figure 1 A block diagram of an example of a streaming media server in FIG.
[0154] Reference Fig.14 , the streaming server 106a may include a processor 420, a memory 430, an operation server 440, a buffer 435, and a training server 480a. Fig.14 The streaming server 106a differs from the streaming server 106 in that the training server 480a stores the neural network 485a instead of the neural network 485.
[0155] The neural network 485a can estimate the streaming media data SRDT and can adjust the resolution by using two inference models including an interactive frame prediction model IFPM and a super-resolution model SRM. The processor 420 trains the interactive frame prediction model IFPM and the super-resolution model SRM, and when the training of the interactive frame prediction model IFPM and the super-resolution model SRM is completed, the interactive frame prediction model IFPM and the super-resolution model SRM can be sent to the client device 101 as a trained merged inference model TIM.
[0156] The interactive frame prediction model IFPM performs frame prediction on the streaming media data SRDT having a low resolution to provide the encoder 450 with a prediction frame PFR having a low resolution, and the encoder 450 encodes the streaming media data SRDT having a low resolution by selectively referring to the prediction frame PFR having a low resolution, thereby increasing or improving the encoding speed. The client device 101 receives the merged inference model TIM, and converts the restored streaming media data having a low resolution into restored streaming media data having a high resolution by using the super-resolution model SRM.
[0157] Fig.15 A streaming media system according to an example embodiment is shown.
[0158] Reference Fig.15 , the streaming media system 100b may include a streaming media server 106b and a client device 101b. In some example embodiments, the streaming media system 100b may also include a repository server 490.
[0159] The streaming media server 106 b may include a processor 420 , a memory 430 , an operating server 440 , and a streaming media card 530 . The processor 420 , the memory 430 , the operating server 440 , and the streaming media card 530 may be operably coupled to each other via a bus 410 .
[0160] The operation server 4000 may include a GPU 445, and each operation of the processor 420, the memory 430, and the operation server 440 may be similar to the reference Figure 6 The description is basically the same.
[0161] The streaming media card 530 may include an encoder 531 such as a CODEC, a processing unit (PU) 532, and a communication interface 533 such as a network interface card (NIC). The encoder 531 and the processing unit 532 may be manufactured as one chip. In some example embodiments, the NIC may include an Ethernet. In addition, the NIC may be referred to as a local area network (LAN) and a device for connecting the streaming media server 106 b to a network.
[0162] The processing unit 532 may receive a plurality of interactive frame prediction models IFPM1, IFPM2, and IFPM3, may select one of the plurality of interactive frame prediction models IFPM1, IFPM2, and IFPM3 as a target interactive frame prediction model, may generate a prediction frame about an object frame of the streaming media data SRDT by applying the streaming media data SRDT, user input, and metadata to the target interactive frame prediction model, and may provide the prediction frame to the encoder 531. The processing unit 532 may send information about the target interactive frame prediction model to the client device 101b through the communication interface 533 as a model synchronization protocol MSP.
[0163] The encoder 531 may encode frames of the streaming media data SRDT by selectively referring to prediction frames to generate encoded streaming media data ESRDT, and may send the encoded streaming media data ESRDT to the client device 101b as a model synchronization protocol MSP through the communication interface 533 .
[0164] The client device 101b may include a streaming application processor 121, a memory 130, an I / O interface 150, a display 160, and a communication interface 170. The streaming application processor 121, the memory 130, the I / O interface 150, the display 160, and the communication interface 170 may be coupled to each other via a bus 110. The streaming application processor 121 may be referred to as an application processor.
[0165] Each operation of the memory 130, the I / O interface 150, and the display 160 may be performed with reference to Figure 7 The description is basically the same.
[0166] The streaming application processor 121 may include a modem 122, a decoder 123 such as a multi-function codec (MFC), and a neural processing unit (NPU) 124. The modem 122 may receive the encoded streaming data ESRDT and the model synchronization protocol MSP through the streaming server 106b.
[0167] The memory 130 may store the interactive frame prediction models IFPM2 and IFPM3 and may provide the interactive frame prediction models IFPM2 and IFPM3 to the NPU 124 .
[0168] The NPU can select a target interactive frame prediction model among the interactive frame prediction models IFPM2 and IFPM3 selected by the streaming server 106b based on the model synchronization protocol MSP, obtain a predicted frame by applying the frame of the user input UIN, metadata and encoded streaming data ESRDT to the target interactive frame prediction model, and can provide the predicted frame to the decoder 123.
[0169] The decoder 123 may decode the encoded streaming data ESDRT by selectively referring to the prediction frame to generate restored streaming data RSRDT, and may provide the restored streaming data RSRDT to the user through the display 160 .
[0170] exist Fig.15 In the embodiment, each of the streaming media card 530 and the streaming media application processor 121 can be implemented by hardware such as a logic circuit, a processing circuit, etc. The streaming media card 530 can be installed on the streaming media server 101, and the streaming media application processor 121 can be installed on the client device 101b. In some example embodiments, when the streaming media card 530 is installed in a personal computer, the personal computer can be used as a streaming media server.
[0171] The repository server 490 may include interactive frame prediction models IFPM1, IFPM2, and IFPM3, may train the interactive frame prediction models IFPM1, IFPM2, and IFPM3, and may send at least some of the interactive frame prediction models IFPM1, IFPM2, and IFPM3 to the training server 106b and the client device 101b upon completion of training of the interactive frame prediction models IFPM1, IFPM2, and IFPM3.
[0172] Fig.16 is a diagram showing a method according to an example embodiment Fig.15 A block diagram of an example of a streaming media card in FIG.
[0173] exist Fig.16 For ease of explanation, a GPU 445, a processor 420, and a memory 430 are also shown.
[0174] Reference Fig.16 , the streaming media card 530a may include a first processing cluster 540, a second processing cluster 550, a first encoder 531a, a second encoder 531b, a first communication interface 533a and a second communication interface 533b. The first communication interface 533a and the second communication interface 533b may both be implemented using NIC.
[0175] GPU 445 may generate first streaming data SRDT1 associated with a first user and second streaming data SRDT2 associated with a second user different from the first user, and may provide the first streaming data SRDT1 and the second streaming data SRDT2 to the first processing cluster 540 and the second processing cluster 550, respectively.
[0176] The first processing cluster 540 can generate a first prediction frame PFR1 by applying the first streaming media data SRDT1 to the first interactive frame prediction model among the multiple interactive frame prediction models, and can provide the first prediction frame PFR1 to the first encoder 531a. The first processing cluster 540 may include: multiple NPUs 541, 543, and 545 with a pipeline configuration; multiple cache memories (CACHE) 542, 544, and 546 connected to the NPUs 541, 543, and 545, respectively; and a spare NPU 547. NPUs 541, 543, and 545 can respectively implement different reasoning models using different neural networks. The spare NPU 547 can adopt a neural network model that will be used later. Cache memories 542, 544, and 546 can all store commonly used data in the corresponding NPUs in NPUs 541, 543, and 545, and can enhance performance.
[0177] The second processing cluster 550 can generate a second prediction frame PFR2 by applying the second streaming media data SRDT2 to a second interactive frame prediction model among multiple interactive frame prediction models, and can provide the second prediction frame PFR2 to the second encoder 531b. The second processing cluster 550 may include: multiple NPUs 551, 553, and 555 with a pipeline configuration; multiple cache memories 552, 554, and 556 connected to the NPUs 551, 553, and 555, respectively; and a spare NPU 557. NPUs 551, 553, and 555 can respectively implement different reasoning models using different neural networks. The spare NPU 557 can adopt a neural network model to be used later. Cache memories 552, 554, and 556 can all store commonly used data in the corresponding NPUs in NPUs 551, 553, and 555, and can enhance performance.
[0178] The first encoder 531a may encode the first streaming data SRDT1 by selectively referring to the first prediction frame PFR1 to generate first encoded streaming data ESRDT1, and may send the first encoded streaming data ESRDT1 to a first client device used by a first user through the first communication interface 533a.
[0179] The second encoder 531b may encode the second streaming data SRDT2 by selectively referring to the second prediction frame PFR2 to generate second encoded streaming data ESRDT2, and may send the second encoded streaming data ESRDT2 to a second client device used by a second user through the second communication interface 533b.
[0180] The first processing cluster 540 and the second processing cluster 550 may be incorporated into Fig.15 In the processing unit 532, the first encoder 531a and the second encoder 531b may be incorporated into Fig.15 In the encoder 531, the first communication interface 533a and the second communication interface 533b may be incorporated into Fig.15 In the communication interface 533.
[0181] The first processing cluster 540 may be Fig.15 The repository server 490 receives information MID1 about the first interactive frame prediction model, and the second processing cluster 550 can obtain information MID1 from Fig.15 The repository server 490 receives information MID2 about the second interactive frame prediction model.
[0182] Fig.17 is a flow chart illustrating the operation of a client device according to an example embodiment.
[0183] Fig.18A and Fig.18B Example operations of a client device are respectively shown.
[0184] exist Figures 17 to 18B In the above example, it is assumed that the interactive frame prediction model is further applied with Fig.14 The super-resolution model in SRM is a model of the resolution adjustment model.
[0185] Reference Figure 6 , Figure 7 , Fig.12 and Figures 17 to 18B , the client device 101 may receive first streaming media data from the streaming media server 106 through the communication interface 170, the first streaming media data corresponding to a first resolution of the original streaming media data associated with the first domain (operation S610). In some example embodiments, the user input may include a user input for selecting a network delay associated with the first domain, or a user input for selecting a first resolution of the original streaming media data associated with the first domain. In some example embodiments, the client device 101 may receive second streaming media data corresponding to a second resolution from the streaming media server 106 based on obtaining the user input during the receipt of the first streaming media data corresponding to the first resolution from the streaming media server 106.
[0186] The client device 101 may select, based on user input, a first interactive frame prediction model corresponding to a first resolution of the original streaming media data from among a plurality of interactive frame prediction models corresponding to a plurality of resolutions of the original streaming media data (operation S620). In some example embodiments, the client device 101 may select at least one interactive frame prediction model SIFPM that meets an image error rate (ER) (e.g., 3%) selected by the user from among a plurality of interactive frame prediction models corresponding to a plurality of resolutions of the original streaming media data. Fig.18A For example, the client device 101 can select at least one interactive frame prediction model that meets the image error rates selected by the user at the first time point T1 and the second time point T2 from a first interactive frame prediction model IFPM11 710 corresponding to a first resolution of the original streaming data associated with the first domain 701 (for example, a resolution of 80% relative to the original resolution), a second interactive frame prediction model IFPM12 720 corresponding to a second resolution of the original streaming data associated with the first domain 701 (for example, a resolution of 60% relative to the original resolution), a third interactive frame prediction model IFPM13 730 corresponding to a third resolution of the original streaming data associated with the first domain 701 (for example, a resolution of 40% relative to the original resolution), and a fourth interactive frame prediction model IFPM14 740 corresponding to a fourth resolution of the original streaming data associated with the first domain 701 (for example, a resolution of 20% relative to the original resolution).
[0187] The client device 101 may decode the first streaming media data into restored first streaming media data using the selected first interactive frame prediction model (operation S630 ). The client device 101 may display the restored first streaming media data on the display 160 .
[0188] Reference Fig.18BFor example, there are four image error rates ER1 751, ER2 752, ER3 761, and ER4 762. The client device 101 may select at least one interactive frame prediction model associated with a specific domain to conform to the image error rate corresponding to a specific resolution selected by the user. The client device 101 may select an interactive frame prediction model IFPM23 that conforms to the image error rate of 5% selected by the user from among the interactive frame prediction models IFPM21, IFPM22, IFPM23, and IFPM24 associated with the first domain 701, and may select an interactive frame prediction model IFPM21 that conforms to the image error rate of 3%. The client device 301 may select an interactive frame prediction model IFPM33 that conforms to the image error rate of 5% selected by the user from among the interactive frame prediction models IFPM31, IFPM32, IFPM33, and IFPM34 associated with the second domain 703, and may select an interactive frame prediction model IFPM34 that conforms to the image error rate of 10%. The selected interactive frame prediction mode may correspond to the lowest resolution of the original streaming media data.
[0189] Fig.19 Example operations of a client device according to example embodiments are shown.
[0190] Reference Fig.19 Referring to reference numeral 901 in the figure, a plurality of interactive frame prediction models corresponding to a plurality of resolutions of the original streaming media data are stored in the training server 106 in an untrained state, and the image error rate ER of each interactive frame prediction model may be "1". The client device 101 may identify a resolution RR 931 associated with a predetermined or optionally desired condition 921 (for example, the original resolution is 100%) from a value 911 obtained by applying the image error rate ER corresponding to the plurality of resolutions to the compression rate CR corresponding to the plurality of resolutions. In the graph indicated by reference numeral 901, since the resolution 931 corresponding to the original resolution of 100% has an image error rate ER of "0", a value associated with the predetermined or optionally desired condition 921 may be obtained at the resolution 931.
[0191] Reference Fig.19In the reference numeral 902 in the figure, a plurality of interactive frame prediction models corresponding to a plurality of resolutions of the original streaming media data are stored in the training server 106 in a trained state after a period of time, and the image error rate ER of each interactive frame prediction model may be different. The client device 101 may identify a resolution RR 932 associated with a predetermined or optionally desired condition 922 (e.g., the original resolution is 80%) from a value 912 obtained by applying the image error rate ER corresponding to the plurality of resolutions to the compression rate CR corresponding to the plurality of resolutions, and may select an interactive frame prediction model corresponding to the resolution 932.
[0192] Fig. 20 Example operations of a client device according to example embodiments are shown.
[0193] Reference Fig. 20 The processor 120 in the client device 101 can identify the network bandwidth NTBW for the client device 101 based on a predetermined or optionally desired time period, a user's request, and a request from the streaming server 106 .
[0194] The client device 101 may send the identified network bandwidth NTBW to the streaming server 106 .
[0195] The streaming server 106 may select an interactive frame prediction model corresponding to the selected resolution from among the multiple interactive frame prediction models IFPM41, IFPM42, and IFPM43 corresponding to the multiple resolutions of the original streaming data based on the compression rates corresponding to the multiple resolutions of the original streaming data, the image error rates corresponding to the multiple resolutions, and the network bandwidth NTBW of the client device 101. The streaming server 106 may send the selected interactive frame prediction model SIFPM corresponding to the identified resolution to the client device 101.
[0196] exist Fig.19 and Fig. 20 In the above example, it is assumed that the interactive frame prediction model is further applied with Fig.14 The super-resolution model in SRM is a model of the resolution adjustment model.
[0197] Fig.21 A training operation of a streaming media server according to an example embodiment is shown.
[0198] Reference Fig.21 The streaming server 106 may generate a plurality of streaming data corresponding to a plurality of resolutions of the original streaming data associated with the plurality of domains.
[0199] For example, the streaming server 106 may generate a plurality of streaming data corresponding to a plurality of resolutions of the original streaming data associated with the first domain 1301. The streaming server 106 may generate a plurality of streaming data corresponding to a plurality of resolutions of the original streaming data associated with the second domain 1302. The streaming server 106 may generate a plurality of streaming data corresponding to a plurality of resolutions of the original streaming data associated with the third domain 1303.
[0200] The streaming server 106 may train multiple interactive frame prediction models corresponding to multiple resolutions of the original streaming data by using the original streaming data associated with a specific domain and the multiple streaming data corresponding to multiple resolutions of the original streaming data.
[0201] For example, the streaming server 106 can train the first interactive frame prediction model IFPM51 by providing the first streaming data corresponding to the first resolution associated with the first domain 1301 as training data, and the original streaming data corresponding to the first resolution associated with the first domain 1301 as expected data to the first interactive frame prediction model IFPM51.
[0202] For example, the streaming server 106 can train the second interactive frame prediction model IFPM52 by providing the second interactive frame prediction model IFPM52 with second streaming data corresponding to the second resolution associated with the second domain 1302 as training data, and original streaming data corresponding to the second resolution associated with the second domain 1302 as expected data.
[0203] For example, the streaming server 106 can train the third interactive frame prediction model IFPM53 by providing the third interactive frame prediction model IFPM53 with third streaming data corresponding to the third resolution associated with the third domain 1303 as training data, and original streaming data corresponding to the third resolution associated with the third domain 1303 as expected data.
[0204] The streaming server 106 may provide the trained interactive frame prediction models IFPM51 , IFPM52 , and IFPM53 to the client device 101 .
[0205] exist Fig.21 In the above example, it is assumed that the interactive frame prediction model is further applied with Fig.14 The super-resolution model in SRM is a model of the resolution adjustment model.
[0206] Fig. 22 is a block diagram illustrating an electronic system according to example embodiments.
[0207] Reference Fig. 22 , the electronic system 1400 includes a video source (SRC) 1410 and a video codec (CODEC) 1420. The electronic system 1400 may further include a processor 1430, a connection module 1440, a storage device 1450, an I / O device 1460, and a power supply 1470.
[0208] The video source 1410 provides the encoded streaming media data ESRDT and the interactive frame prediction model IFPM. For example, the video source 1410 may include a streaming media server for providing a streaming media service. According to an example embodiment, the video source 1410 may include an encoder. The encoder may encode the streaming media data by selectively referring to the prediction frame provided from the interactive frame prediction model IFPM to provide the encoded streaming media data ESRDT.
[0209] According to example embodiments, the video codec 1420 may include a decoder.
[0210] The processor 1430 may perform various computing functions, such as specific calculations and tasks. The connection module 1040 may communicate with an external device and may include a transmitter 1442 and / or a receiver 1444. The storage device 1450 may operate as a data storage for data processed by the electronic system 1400, or as a working memory.
[0211] The I / O device 1460 may include at least one input device such as a keypad, buttons, a microphone, a touch screen, etc., and / or at least one output device such as a speaker, a display device 1062 , etc. The power supply 1470 may provide power to the electronic system 1000 .
[0212] Aspects of the inventive concept may be applied to various streaming media servers that provide streaming media services.
[0213] The foregoing is illustrative of example embodiments and should not be construed as limiting thereof. Although some example embodiments have been described, those skilled in the art will readily appreciate that many modifications may be made in the example embodiments without materially departing from the novel teachings and advantages of the inventive concept. Therefore, all such modifications are intended to be included within the scope of the inventive concept as defined in the claims.
Claims
1. A streaming media system, comprising: A streaming media server, wherein the streaming media server is configured to: training an interactive frame prediction model based on the streaming media data, the user input, and metadata associated with the user input, encoding the streaming media data by selectively using a predicted frame generated based on the trained interactive frame prediction model, wherein the trained interactive frame prediction model receives as input a user input, streaming media data and metadata associated with the user input and generates as output the predicted frame, and Sending the trained interactive frame prediction model and the encoded streaming media data; and A client device, wherein the client device is configured to: receiving the trained interactive frame prediction model and the encoded streaming media data, and The encoded streaming media data is decoded based on the trained interactive frame prediction model to provide the restored streaming media data to the user.
2. The streaming media system according to claim 1, wherein: The streaming media server comprises: processor; a memory, the memory being coupled to the processor, the memory storing instructions; an operation server, the operation server being coupled to the processor, the operation server comprising an encoder and a graphics processing unit configured to generate the streaming media data; and A training server is coupled to the processor, the training server being configured to store a neural network configured to implement the interactive frame prediction model.
3. The streaming media system according to claim 2, wherein: The processor is configured to execute the instructions so that: The graphics processing unit is configured to provide the streaming media data to the training server; and The training server is configured as follows: applying the frames of the streaming media data, the user input, and the metadata to the interactive frame prediction model to train the interactive frame prediction model so that the interactive frame prediction model provides a predicted frame about an object frame of the streaming media data, When the training of the interactive frame prediction model is completed, the trained interactive frame prediction model is sent to the client device, and The predicted frame is provided to the encoder.
4. The streaming media system according to claim 3, wherein: The processor is configured to determine that training of the interactive frame prediction model is completed in response to a difference between a compression rate of the predicted frame and a compression rate of an expected frame associated with the predicted frame being within a reference range.
5. The streaming media system according to claim 4, wherein: The processor is configured to execute the instructions so that: The encoder is configured as: Encoding the object frame by referring to a higher similarity frame selected from a previous frame of the streaming media data and the predicted frame, and providing the encoded streaming media data to the client device, the higher similarity frame having a higher similarity with the object frame, the previous frame being a previous frame of the object frame; as well as Motion estimation is performed by referring to the higher similarity frame.
6. The streaming media system according to claim 2, wherein: The training server is configured to adjust the resolution of the frame of the streaming media data by further applying a super-resolution model to the frame of the streaming media data.
7. The streaming media system according to claim 1, wherein: The client device comprises: monitor; Communication interface; an input / output interface, the input / output interface being configured to receive the user input; a processor coupled to the display, the input / output interface, and the communication interface; and A memory is coupled to the processor, the memory storing instructions.
8. The streaming media system according to claim 7, wherein: The processor is configured to execute the instructions so that: The input / output interface is configured to provide the user input and the metadata based on the user input to the streaming media server through the communication interface; and The communication interface is configured to receive the trained interactive frame prediction model to store the trained interactive frame prediction model in the memory, thereby generating a predicted frame by applying the encoded streaming data, the user input and the metadata to the trained interactive frame prediction model.
9. The streaming media system according to claim 8, wherein: The processor includes a decoder configured to decode the encoded streaming media data by selectively using the prediction frame to generate the restored streaming media data; and The processor is configured to execute the instructions to display the restored streaming media data on the display.
10. The streaming media system according to claim 7, in, The trained interactive frame prediction model corresponds to a model to which a resolution adjustment model is also applied, The processor is configured to execute the instructions so as to: receiving, via the communication interface, first encoded streaming media data corresponding to a first resolution of original streaming media data associated with a first domain; selecting a first trained interactive frame prediction model from among a plurality of trained interactive frame prediction models corresponding to a plurality of resolutions of the original streaming media data, decoding the first encoded streaming media data into first restored streaming media data based on the first trained interactive frame prediction model, and The first restored streaming media data is displayed on the display.
11. The streaming media system according to claim 10, wherein: The first domain includes a plurality of subdomains corresponding to a plurality of areas.
12. The streaming media system according to claim 10, wherein: The processor is configured to select the first trained interactive frame prediction model corresponding to the first resolution based on the user input.
13. The streaming media system according to claim 1, wherein: The client device supports virtual reality.
14. A streaming media system, comprising: A streaming media server, wherein the streaming media server is configured to: Selecting a target interactive frame prediction model among multiple interactive frame prediction models, training the target interactive frame prediction model based on streaming media data, user input, and metadata associated with the user input, generating a predicted frame by using the target interactive frame prediction model, wherein the target interactive frame prediction model receives as input a user input, streaming media data, and metadata associated with the user input and generates as output the predicted frame, encoding the streaming media data by selectively using the predicted frames, and Sending the encoded streaming media data; and A client device, wherein the client device is configured to: receiving the plurality of interactive frame prediction models and the encoded streaming media data, selecting the target interactive frame prediction model among the plurality of interactive frame prediction models, The encoded streaming media data is decoded based on the target interactive frame prediction model to provide the restored streaming media data to the user.
15. The streaming media system according to claim 14, wherein: The streaming media server comprises: processor; a memory, the memory being coupled to the processor, the memory storing instructions; an operation server, the operation server being coupled to the processor, the operation server comprising an encoder and a graphics processing unit configured to generate the streaming media data; and A streaming media card is coupled to the processor, and is configured to generate the encoded streaming media data by using the target interactive frame prediction model.
16. The streaming media system according to claim 15, wherein: The streaming media card comprises: At least one processing unit, the at least one processing unit being configured to: selecting the target interactive frame prediction model among the plurality of interactive frame prediction models, and generating the predicted frame by applying the streaming media data, the user input, and the metadata to the target interactive frame prediction model; at least one encoder configured to generate the encoded streaming media data by encoding an object frame of the streaming media data with reference to a higher similarity frame selected from a previous frame of the streaming media data and the predicted frame, the higher similarity frame having a higher similarity to the object frame, the previous frame being a previous frame of the object frame; and A communication interface is configured to provide the encoded streaming media data to the client device.
17. The streaming media system according to claim 16, wherein: The at least one processing unit comprises: a first processing cluster configured to generate a first prediction frame by applying first streaming media data associated with a first user to a first interactive frame prediction model among the plurality of interactive frame prediction models; and a second processing cluster configured to generate a second prediction frame by applying second streaming media data associated with a second user different from the first user to a second interactive frame prediction model among the plurality of interactive frame prediction models, and The first processing cluster and the second processing cluster support the multiple interactive frame prediction models through a pipeline solution.
18. The streaming media system according to claim 14, wherein: The client device comprises: monitor; Communication interface; an input / output interface, the input / output interface being configured to receive the user input; an application processor coupled to the display, the input / output interface, and the communication interface; and a memory, the memory being coupled to the application processor, the memory storing instructions, The application processor is configured to execute the instructions, thereby: storing the plurality of interactive frame prediction models received via the communication interface in the memory; Applying the encoded streaming media data, the user input and the metadata to the target interactive frame prediction model among the multiple interactive frame prediction models to generate a prediction frame; Decoding the encoded streaming media data based on the predicted frame to generate restored streaming media data; and The restored streaming media data is displayed on the display.
19. The streaming media system according to claim 14, further comprising: A repository server, the repository server being configured to: storing the plurality of interactive frame prediction models, training the plurality of interactive frame prediction models, and The trained multiple interactive frame prediction models are provided to the streaming media server and the client device.
20. A streaming media server, comprising: processor; a memory, the memory being coupled to the processor, the memory storing instructions; an operation server, the operation server being coupled to the processor, the operation server comprising an encoder and a graphics processing unit configured to generate streaming media data; as well as a training server coupled to the processor, the training server configured to store a neural network configured to implement an interactive frame prediction model, The processor is configured to execute instructions such that: The training server is configured to: train the interactive frame prediction model based on the streaming media data, the user input, and metadata associated with the user input; The encoder is configured to encode the streaming media data by selectively using a prediction frame generated based on the trained interactive frame prediction model, wherein the prediction frame is generated and output by the trained interactive frame prediction model by receiving user input, streaming media data, and metadata associated with the user input as input; and The training server is configured to send the trained interactive frame prediction model and the encoded streaming media data to the client device.
Citation Information
Patent Citations
Braising material and manufacturing methods thereof, and braising methods thereby
KR1020200045787A
End-to-end video and image compression
US10623775B1